Saturday, May 16, 2009

The Failure of Search (or the Fallacy of Abundance)

I originally wrote this entry on October 1, 2004, and published it on blogs.sun.com.


So, what's going on with web search? Why is it giving such high rank to (at best) marginal material, such as the one on this weblog on certain topics? Or as I asked earlier, why do we even feel that we get anything relevant when we perform a search on the Web? How much better material are we actually missing when we limit ourselves to the findings of a search engine?


It is in asking those sorts of questions that we can arrive at modest discoveries or at least novel explanations of what we see around us.


To further the investigation I reported earlier, I went back to the chapters on search in Hubert Dreyfus' little book, On the Internet. According to Dreyfus, given the immense size of the Net, it is "estimated that search engines can recall at most 2 per cent of the relevant sites." (The number might have changed in the last three years but I don't believe that the changes, if any, would affect the arguments in any drastic way.)


We need to ask why "content" (or "information") retrieval systems are receiving the hype they are receiving even if they are hardly adequate when it comes to searching for specific content. How could my weblogs, even if they are somewhat useful, be ranked as the third most useful or important content on certain scholars I've only occasionally quoted and on whose works I still consider myself a novice?


Surely, this sort of system behavior cannot be good if we have hopes to be able to find important bits of documents or knowledge through search and information retrieval.


To explain the hype regarding search and information retrieval, Dreyfus quotes computer scientist David Blair, who cites information retrieval (IR) pioneer Don Swanson:




IR prioneer Don Swanson observed this phenomenon decades ago, and calls it the "fallacy of abundance". The fallacy of abundance is the mistake a searcher makes when he uses a large IR system and is able to find some useful documents. Swanson pointed out that on a sufficiently large system . . . almost any query will retrieve some useful documents. The mistake is to think that just because you got some useful documents the IR system is performing well. What you don't know is how many better documents the system missed.



And so . . . since my weblogs can be ranked highly by Google for certain subjects, they may be perceived (by some searchers) to be more important than they really are.



Do Google Rankings Mean Anything

I originally wrote this entry on October 1, 2004, and published it on blogs.sun.com.


Simple errors and break-downs often lead to new discoveries . . . The type of break-down is almost immaterial.


Error:


A few days ago, I accidentally removed all records of referers and hits on my weblog, all 106,500 of them.


This simple push-of-a-button dropped me out of the hot list you see at the bottom of blogs.sun.com.


Curiosity:


Before this incidence, I'd never thought much about the referers list but deleting all the hit records made me curious. So, I have now gone back to the gradually accumulating referers list for this weblog and have tried to learn something about the referer URL distribution. A significant majority of the hits on this weblog are direct and an equally significant minority are from Google (and competing search engine) searches.


Example:


Although I'm not sure how persistent this sort of system behavior is, this weblog is currently (as of early Octobor, 2004) receiving high (Google and other) search rankings on subjects in which the author is barely a novice.


The rankings this weblog is receiving from Google (as well as other competing search engines) for certain specific queries, for example queries on Oliver Williamson (see Ref.1) and on Chester Barnard (see Ref.2) truly amaze me. As of earlier this week, I've consistently been ranked third on both (and their varient orderings) on Google. See Ref.1.1 and Ref.2.1 for the relevant Google searches.


I may have had the good fortune of having studied with Oliver Williamson for a very short period of time but I'm still a novice learner of his ideas. I may have studied portions of Chester Barnard's classical book, because Williamson recommended it, but I do not deserve to be read diligently as serious commentary on either. So, why is it that what I have written about both is receiving high Google rankings. Surely, I myself know better written material on both topics.


Questions:


What's broken down? What's amiss about search, whether of the Google variety or not? Why do we even feel that we get anything relevant when we perform a search on the Web? How much better material are we actually missing if we limit ourselves to the findings of a search engine?


A Modest Discovery:


Web search and information retrieval fails us more often than we know or realize !


Friday, May 08, 2009

In Search of Chester Barnard


I originally wrote this entry on September 29, 2004, and published it on blogs.sun.com.


As Google's stocks surge in price, the WSJ reports:



Indeed, Google shares now are at heady prices, trading at a whopping 55 times next year's expected earnings of $2.29 per share, compared with a price-to-earnings multiple of 15 for the Standard & Poor's 500-stock index. And part of the reason the shares are rallying is that Google has fewer shares that are freely available to trade, compared with comparable companies, so buying interest goes a lot further in pushing the stock higher. Google has a "float" of almost 30 million shares, compared with more than one billion shares of Yahoo that trade freely.




But how good or useful is Internet search?


I certainly use it all the time but I also end up having to filter a great deal of nonsense. There are also cases that produce some amazement. As of last week, Google has been ranking this Sun weblog third among 110,000 finds on Chester Barnard. Another search site (A9, which I believe must be using the same Google technology for search) gives this weblog the same ranking on Chester Barnard.


How good of an expert am I on Chester Barnard? Do my very casual writings on his work really deserve to be of such high ranking in search results? I doubt it very much. I'm just a novice and an amateur reader of Barnard's essential writings on organizational theory. I guess the only thing I've brought into the fold is to connect Barnard's work with others' and to give a few useful URLs to follow, but reading what I've written about his work will not make any one an expert either. For that, a different kind of practice and training would be required. To begin with, one should probably start reading Oliver Williamson's "Chester Barnard and the Incipient Science Of Organization" published in his The Mechanisms of Governance and also in his Organization Theory: From Chester Barnard to the Present and Beyond.



Spyware Vote in the Congress

I originally wrote this entry on September 22, 2004, and published it on blogs.sun.com.


The U.S. House of Representatives is expected to vote next week on a measure to crack down on spyware. (The relevant bills originally took shape in the Commmittee on Energy and Commerce.)


Reuters reports the measure has wide support.


I've not been able to find and read any of the two proposed bills (which are supposed to be merged by the powerful Committee on Rules) but I'm wondering about their scope and how severe the punishement will be for the offenders.

The Smoothest Transition in the World


I originally wrote this entry on September 20, 2004, and published it on blogs.sun.com.


On the PC:


IE to Firefox.