Thursday, January 10, 2008

Joining Microsoft Live Labs

I am starting at Microsoft Live Labs next week.

Live Labs is an applied research group affiliated with Microsoft Research and MSN. The group has the enjoyable goal of not only trying to solve hard problems with broad impact, but also getting useful research work out the door and into products so it can help as many people as possible as quickly as possible.

Live Labs is lead by Gary Flake, the former head of Yahoo Research. It is a fairly new group, formed only two years ago. Gary wrote a manifesto that has more information about Live Labs.

With its connection to MSR, working at Live Labs promises to be a lot of fun. I am thrilled that I might have an opportunity to work with Sue Dumais, Jaime Teevan, Dennis DeCoste, Matt Hurst, Eric Brill, Steve Drucker, Paul Viola, Patrice Simard, and many, many others.

Slides on cluster computing and MapReduce

Google has published slides and video from a course taught to interns during the summer of 2007.

The slides are well worth reading whether you are new to these topics or consider yourself an old pro. They cover introductory topics in distributed systems, motivate and describe MapReduce, and then discuss how to implement the PageRank algorithm and a variant on K-means clustering in parallel on a large cluster using MapReduce.

Thanks, Maxime, for pointing out these slides and videos.

Sunday, January 06, 2008

MapReducing 20 petabytes per day

Googlers Jeff Dean and Sanjay Ghemawat have an article, "MapReduce: Simplified Data Processing on Large Clusters", in the January 2008 issue of Communications of the ACM.

It is a great introduction to MapReduce, but what I found most interesting was the numbers they cite on usage of MapReduce at Google.

Jeff and Sanjay report that, on average, 100k MapReduce jobs are executed every day, processing more than 20 petabytes of data per day.

More than 10k distinct MapReduce programs have been implemented. In the month of September 2007, 11,081 machine years of computation were used for 2.2M MapReduce jobs. On average, these jobs used 400 machines and completed their tasks in 395 seconds.

What is so remarkable about this is how casual it makes large scale data processing. Anyone at Google can write a MapReduce program that uses hundreds or thousands of machines from their cluster. Anyone at Google can process terabytes of data. And they can get their results back in about 10 minutes, so they can iterate on it and try something else if they didn't get what they wanted the first time.

It is an amazing tool for Google. Google's massive cluster and the tools built on top of it rightly have been called a "competitive advantage", the "secret source of Google's power", and a "major force multiplier".

By the way, there is another little tidbit in the paper about Google's machine configuration that might be of interest. They describe the machines as dual processor boxes with gigabit ethernet and 4-8G of memory. The question of how much memory Google has per box in its cluster has come up a few times, including in my previous posts, "Four petabytes in memory?" and "Power, performance, and Google".

Update: Others now have picked up on this paper and commented on it, including Niall Kennedy, Kevin Burton, Paul Kedrosky, Todd Huff, Ionut Alex Chitu, and Erick Schonfeld.

Update: Let me also point out Professor David Patterson's comments on MapReduce in the preceding article in the Jan 2008 issue of CACM. David said, "The beauty of MapReduce is that any programmer can understand it, and its power comes from being able to harness thousands of computers behind that simple interface. When paired with the distributed Google File System to deliver data, programmers can write simple functions that do amazing things."

Friday, January 04, 2008

The coming 2008 dot-com crash

Early January is the time we see many predictions for 2008. I have not played this game since 2006, but I want to chime in this year.

I am only going to make one prediction, but one with broad impact. We will see a dot-com crash in 2008. It will be more prolonged and deeper than the crash of 2000.

The crash will be driven by a recession and prolonged slow growth in the US. Global investment capital will flee to quality, ending the speculative dumping of cash on Web 2.0 startups.

Venture capital firms will seek to limit their losses by forcing many of their portfolio companies to liquidate or seek a buyout. Buyout prospects will be poor, however, as the cash rich companies find themselves in a buyers market and let those seeking a savior come face-to-face with the spectre of bankruptcy before finally buying up the assets on the cheap.

Startups that managed to get cash before the bubble collapses will have a cash horde, but will find little opportunity to rest on it. Most startups will find their revenue models were unrealistic and will rapidly have to seek change. Many will jump over to advertising, but the advertising market will have constricted. Bigger businesses will seek to drive out the new entrants, and online advertising will become a cutthroat business with little profits to be found. Others startups may shift toward licensing and development deals for bigger companies, but will find their investors impatient now that the promised $500M startup has become a $10M company.

The big players will not be immune from this contagion. Google, in particular, will find its one-trick pony lame, with the advertising market suddenly stagnant or contracting and substantial new competition. The desperate competition with dwindling opportunity will drive profits in online advertising to near zero. Google and Yahoo will find their available cash dropping and will do substantial layoffs.

Unfortunately, this scenario has privacy implications as well. Much like we saw after the 2000 crash, it is likely that those with little to lose will attempt scary new forms of advertising. The Web will become polluted with spyware, intrusiveness, and horrible annoyances. None of this will work, of course, and there will be lawsuits and new privacy legislation, but we will have to endure it while it lasts.

It is a dire scenario, but one that looks much like what we saw after 2000. That was a much smaller crash without the fuel from broader problems in the US economy, but we still had investment capital shut off for a few years, most startups shut down, and the remaining startups shift business models. We also saw a dramatic rise in pop-up advertising and spyware.

The crash of 2008 will be similar to 2000 but deeper. We all will have to weather the storm.

Thursday, January 03, 2008

A brief history of Findory

Findory logoFindory was a personalized news site. The site launched in January 2004 and shut down November 2007.

A reader first coming to Findory would see a normal front page of news, the popular and important news stories of the day. When someone read articles on the site, Findory learned what stories interested that reader and changed the news that was featured to match that reader's interests. In this way, Findory built each reader a personalized front page of news.

Below is a screenshot of an example personalized Findory home page. Articles marked with a sunburst icon are personalized, picked specifically for this reader based on this person's reading history.

[Clicking on the screenshot will bring up a full-sized version]

Findory's personalization used a type of hybrid collaborative filtering algorithm that recommended articles based on a combination of similarity of content and articles that tended to interested other Findory users with similar tastes.

One way to think of this is that, when a person found and read an interesting article on Findory, that article would be shared with any other Findory readers who likely would be interested. Likewise, that person would benefit from interesting articles other Findory readers found. All this sharing of articles was done implicitly and anonymously without any effort from readers by Findory's recommendation engine.

Findory's news recommendations were unusual in that they were primarily based on user behavior (what articles other readers had found), worked from very little data (starting after a single click on Findory), worked in real-time (changed immediately when someone read an article), required no set-up or configuration (worked just by watching articles read), and did not readers to identify themselves (no login necessary).

Findory's primary product was in news, but the broader goal of Findory was to personalize information. Toward that, Findory had alpha features that would recommend videos, podcasts, feeds, advertisements, and web search results.

Video, podcast, and feed recommendations worked much like the news recommendations. The advertisement recommendations were an unusual form of fine-grained personalized advertising that attempted to target advertisements based not only on the content of the page, but also a person's reading history on Findory. The web search was an unusual form of fine-grained personalized web search that modified Google search results to feature items clicked on by searchers with similar search behavior (recommendations) or that were clicked on by this specific searcher in the past (re-finding).

At its height, Findory was a popular website with over 100k unique visitors and 5M page views per month. Findory was well reviewed and received press coverage in the Wall Street Journal, Forbes, Time Magazine, PC World, The Times, Spiegel, Seattle PI, Seattle Times, Puget Sound Business Journal, KPLU, Slate, and elsewhere.

More information and more details on Findory's history can be found in my many previous posts on Findory.

Wednesday, January 02, 2008

Upcoming Yahoo talk on computational advertising

Andrei Broder from Yahoo Research will be giving a talk on "Computational Advertising" next week (Thursday, Jan 10) at University of Washington. Video of the talk will be available live and archived a few days afterward.

The talk looks like a great one. Excerpts from the description:
Computational advertising ... [attempts to] find the "best match" between a given user in a given context and a suitable advertisement.

The information about the user can vary from scarily detailed to practically nil. The number of potential advertisements might be in the billions. Thus, depending on the definition of "best match" this challenge leads to a variety of massive optimization and search problems, with complicated constraints.

This talk will give an introduction to this area and give a "taste" of some recent results.
From the way Andrei is framing the problem -- matching advertisements not only to content, but also to what we know about each user -- Andrei clearly is talking about personalized advertising.

Personalized advertising is a tremendous computational challenge. Traditional contextual advertising matches ads to static content. We only have to do the match infrequently, then we can show a selection of the ads that we think will work well for a given piece of content to everyone who views that content.

With personalized advertising, we match ads to content and each user's interests, and then show different ads for each user. Like with all personalization, caching no longer works. Each user sees a different page. With personalized advertising, targeting ads now means we have to find matches in real-time for each page view and each user.

On a related note, Yahoo Researcher Omid Madani and ex-Yahoo Researcher Dennis DeCoste had an interesting short paper back in 2005, "Contextual Recommender Problems" (PDF), that has some more thoughts on this problem. As I wrote in an older post, Omid and Dennis treated personalized advertising as a recommendations problem and proposed a few methods of attacking the problem.

By the way, this talk seems to be a shift for Andrei away from a more general problem of personalized information -- which he called "information supply" -- toward focusing on the more specific task of personalized advertising.

See also the Computational Advertising page at Yahoo Research and its list of team members and papers.

Update: It appears this talk will not be archived. It will be broadcast live. If you want to see it remotely, you will have to watch it at the time of the talk.

Update: It was an interesting talk, but, frankly, a bit disappointing in its lack of depth.

Andrei spent most of the talk describing the state of online advertising today, including market size and how targeted advertising works. He touched on some of the more interesting and harder problems, but only touched on them, and only very briefly.

For example, on one slide, Andrei criticized Google AdSense for showing ads for Libby shoes on an article about Dick Cheney and Scooter Libby, saying that the match is spurious. But, Andrei did not say what would be a better ad to show for that news article. In response to my question later, Andrei did say that perhaps no ad is appropriate in that case, but he did not expand on this to talk about how to detect, in general, when it might be undesirable to show ads because of lack of value and commercial intent. When I did a follow-up question after the talk, he expanded briefly into ideas around personalized advertising -- showing ads that might interest this user based on this person's history rather than ads targeting the current content -- and an advertising engine that explores and attempts to learn what ads might be effective, but not in any depth.

For another example, on one slide, Andrei drew a parallel between web search and advertising search, arguing that both can be seen as searches for information, but pointed out that advertisements are a smaller database of smaller documents and that the relevance rank of a search for ads depends on the bids. He did not discuss the issue that web search in some ways is an easier problem, though, in that the results are more easily cached. Web search relevance rank is static over substantial periods of time, but advertisements are not because the relevance of ads depends not only on keyword matching, but also on bids, competing bids, budgets, and clickthrough rates, all of which can vary rapidly.

For a third example, Andrei briefly mentioned using a user profile for personalized advertising, but only touched the surface of what that profile should contain, how it should be used, where it should be stored, in what cases personalized advertising is likely to outperform unpersonalized advertising, and how trying to show different ads to different people massively increases the computation necessary for ad targeting.

The details that did come on the hard problems mostly were in the form of references off to other papers. When talking about how to approximately match keywords picked for ads to the keywords for content, Andrei mentioned Ribeiro-Neto et al. "Impedance coupling in content-targeted advertising" and Yih et al., "Finding Advertising Keywords on Web Pages" (PDF), the latter of which is excellent, by the way. When talking about determining intent, Andrei referred to two of his own recent papers, "A semantic approach to contextual advertising" and "Robust classification of rare queries using web knowledge".

Overall, there was an unfortunate lack of detail on how to solve the "best match" massive optimization challenge of online advertising under all its complicated constraints. Andrei did hint at one point that all the search giants are reluctant to talk about these details, but it is too bad that we were not be able to explore the fun issues in more depth.

Update: Andrei gave a version of his talk at WWW 2008. During the question and answer time, I asked him to expand on what is the "best match" for an advertisement given a user and a context.

He described three major categories of utility: advertiser utility, user utility, and publisher utility. In response to further questions, he suggested that advertiser utility is complicated by the fact that ad agencies may have different incentives than the advertiser and by branding effects. He also pointed out that publisher utility is not as simple as just revenue because of publisher branding issues (e.g. the New York Times will not accept ads for pornography).

As for user utility, suggested that it is complicated and difficult to measure, but that clickthrough rate may be one proxy for it.

Andrei did not expand on how these utility functions could or should be combined or how to deal with conflicts between them.

Questions for 2007 from the NYT

The NYT Bits blog has some fun "Questions We Thought, But Didn't Ask, in 2007" ([1] [2]).

My favorites, first on Web advertising:
I am married with a house. Why do I see so many ads for online dating sites and cheap mortgages?

Should I be happy that I see those ads? It means Internet advertisers still have no idea who I am.
Then on Facebook:
As my number of Facebook friends inevitably expands, with second and third-tier acquaintances and complete strangers joining my network (I am too nice to deny them), doesn't the value of my "social graph" decline?

If Facebook users routinely say they ignore the ads on the site, how has the company become so valuable?

If Internet supremacy is inherently ephemeral ... why isn't their inevitable declines baked into the stratospheric valuations of today's online leaders?
On a related note, Saul Hansell has some "New Questions for a New Year" with some harsh thoughts on the search giants, including Google's inability to "create a significant advertising business for any format other than text ads", Yahoo's failure to become "the best company to work for" which is leaving them with nothing but "a site that is just an old habit in need of changing", and Microsoft's need to go "through MSN and Windows Live with an honest assessment of their business prospects" and determine "how many of the new initiatives" are "rational" investments. As for MySpace, Saul snipes, "Does anyone care anymore?"

Monday, December 24, 2007

Papers from WSDM 2008 on click position bias and social bookmark data

The WSDM conference is being held Feb 11-12 at Stanford University. I am not sure I will make it down from Seattle for it, but, if you are in the SF Bay Area and are interested in search and data mining on the Web, it is an easy one to attend.

Most of the papers for the conference do not appear to be publicly available yet, but, of the ones I could find, I wanted to highlight two of them.

Microsoft Researchers Nick Craswell, Onno Zoeter, Michael Taylor and Bill Ramsey wrote "An Experimental Comparison of Click Position-Bias Models" (PDF) for WSDM 2008. The work looks at models for how "the probability of click is influenced by ... [the] position in the results page".

The basic problem here is that just putting a search result high on the page tends to get it more clicks even if that search result is less relevant than ones below it. If you are trying to learn which results are relevant by looking at which ones get the most clicks, you need to model and then attempt to remove the position bias.

The authors conclude that a "cascade model" which assumes "that the user views search results from top to bottom, deciding whether to click each result before moving to the next" most closely fits searcher click behavior when they look at the top of the search results. However, their "baseline model" -- which assumes "users look at all results and consider each on its merits, then decide which results to click" (that is, position does not matter) -- seemed most accurate for items lower in the search results.

The authors say this suggests there may be "two modes of results viewing", one where searchers click the first thing that looks relevant in the top results, but, if they fail to find anything good, they then shift to scanning all the results before clicking anything.

By the way, if you like this paper, don't miss Radlinski & Joachims' work on learning relevance rank from clickstream data. It not only discusses positional bias in click behavior in search results, but also attempts the next and much more ambitious step of optimizing relevance rank by learning from click behavior. The Craswell et al. WSDM 2008 paper does cite some older work by Joachims and Radlinski, but not this fun and more recent KDD 2007 paper.

The second paper I wanted to point out is by Paul Heymann, Georgia Koutrika and Hector Garcia-Molina at Stanford, "Can Social Bookmarks Improve Web Search?" I was not able to find that paper, but a slightly older tech report (PDF) with the same title is available (found via ResourceShelf). This paper looks at whether data from social bookmark sites like del.icio.us can help us improve Web search.

This is a question that has been subject to much speculation over the last few years. On the one hand, social bookmark data may be high quality labels on web pages because they tend to be created by people for their own use (to help with re-finding). On the other hand, manually labeling the Web is a gargantuan task and it is unclear if the data is substantially different than what we can extract automatically.

Unfortunately, as promising as social bookmarking data might seem, the authors conclude that it is not likely to be useful for Web search. While they generally find the data to be of high quality, they say the data only covers about 0.1% of the Web, only a small fraction of those are not already crawled by search engines, and the tags in social bookmarking data almost always are "obvious in context" and "would be discovered by a search engine." Because of this, the social bookmarking data "are unlikely to be numerous enough to impact the crawl ordering of a major search engine, and the tags produced are unlikely to be much more useful than a full text search emphasizing page titles."

Update: It looks like I will be attending the WSDM 2008 conference after all. If you are going, please say hello if you see me!

Friday, December 21, 2007

Interactive machine learning talk

Dan Olsen at BYU gave a talk, "Interactive Machine Learning", at UW CS a couple months back.

Dan's group is doing some clever work that combines machine learning and HCI. The UW CS talk is good but long. If you are short on time, first take a look at the fun short demo videos Dan's group produced.

I particularly recommend seeing the clever "Screen Crayons" (WMV) application for annotating documents and the "Teaching Robots to Drive" (WMV) demo of an intuitive interface for training a robot car. The "Image Processing with Crayons" (WMV) demo is also good for getting quick introduction to the core idea.

Tuesday, December 18, 2007

Findory turns off the lights

Findory turned off its last webserver today. Sadness.

Previous posts ([1] [2] [3] [4] [5] [6]) have more details on the shutdown and Findory's history.

Monday, December 17, 2007

Microsoft and intelligent agents on the desktop

John Markoff quotes Microsoft Chief Research Officer Craig Mundie on the opportunity multicore processors on our desktop creates for AI and personalized software agents:
In the future, Mr. Mundie said, parallel software will take on tasks that make the computer increasingly act as an intelligent personal assistant.

"My machine overnight could process my in-box, analyze which ones were probably the most important, but it could go a step further," he said. "It could interpret some of them, it could look at whether I've ever corresponded with these people, it could determine the semantic context, it could draft three possible replies. And when I came in in the morning, it would say, hey, I looked at these messages, these are the ones you probably care about, you probably want to do this for these guys, and just click yes and I'll finish the appointment."
Craig had more extensive thoughts at the July 2007 Microsoft Analyst meeting on personalized assistants running on your desktop PC.

Eric Enge interviews Sep Kamvar

Eric Enge posted an interview with Google personalization guru Sep Kamvar.

Some highlights of Sep's answers below:
The two signals that we use right now are the search history and the location. We constantly experiment with other signals, but the two signals that have worked best for us are location and search history.

Some signals that you expect would be good signals, turn out not to be that good. So for example, we did one experiment with Orkut, and we tried to personalize search results based on the community that users had joined. It turns out that while people were interested in the Orkut communities, they didn't necessarily search in line with those Orkut communities.

It actually harkened back to another experiment that we did, where in our first data launch of personalized search we allowed everybody to just check off categories that were of interest to them. People did that, and people would check off categories like literature. Well, they were interested in literature, but they actually didn't do any searching in literature. So, what we thought would be a very clean signal, actually turned out to be a noisy signal.

When I think about what I am interested in, I don't necessarily think about what I am interested in that I search for and what I am interested in that I don't search for. That's something that we found was better learned algorithmically rather than directly.

A signal should be very closely aligned with search and what you are searching for in order for it to be useful to personalizing search ... In addition, we've found that your more recent searches are much more important than searches from a long time ago.
For more on the problems with explicitly extracting preferences -- as Sep found when explicitly asking for each user's category interests in an early version of Google Personalized Search -- please see my post, "Explicit vs. implicit data for news personalization", and the links from that post.

For more on trying to use signals not closely aligned with search, please see my earlier post, "Personalizing search using your desktop files".

For more on the importance of focusing on recent searches for personalized search, please see also my past posts, "The effectiveness of personalized search" and "The many paths of personalization".

Thursday, December 13, 2007

Geoffrey Hinton on the next generation of NNets

AI guru Geoffrey Hinton recently gave a brilliant Google engEdu talk, "The Next Generation of Neural Networks".

If you have any interest in neural networks (or, like me, got frustrated and lost all interest in the mid-1990s), set aside an hour and watch the talk. It is well worth it.

The talk starts with a short description of the history of neural networks, focusing on the frustrations encountered, and then presents Boltzmann machines as a solution.

Geoffrey clearly is motivated by trying to imitate the "model the brain could be using." For example, after fixing the output of a model to ask it to "think" of the digit 2, he enthusiastically describes the model as his "baby", the internal activity of one of the models as its "brain state", and the output of different forms of digits it recognizes as a 2 as "what is going on in its mind."

The talk is also full of enjoyably opinionated lines, such as when Geoffrey introduces alternating Gibbs sampling as a learning method and says, "I figured out how to make this algorithm go 100,000 times faster," adding, with a wink, "The way you do it is instead of running for [many] steps, you run for one step." Or when he dismissively calls support vector machines "a very clever type of perceptron." Or when he criticizes locality sensitive hashing as being "50 times slower [with] ... worse precision-recall curves" than a model he built for finding similar documents. Or when he said, "I have a very good Dutch student who has the property that he doesn't believe a word I say" when talking about how his group is verifying his claim about the number of hidden unit layers that works best.

You really should set aside an hour and watch the whole thing but, if you can't spend that much time or can't take the level of detail, don't miss the history of NNets in the first couple minutes, the description and demo of one of the digit recognition models starting at 18:00 (slide 19), and the discussion of finding related documents using these models starting at 31:40 (slide 28).

On document similarity, it was interesting that a Googler asked a question about finding similar news articles in the Q&A at the end of the talk. The question was about dealing with substantial changes in the types of documents you see -- big news events, presumably, that cause a bunch of new kinds of articles to enter the system -- and Geoffrey addressed it by saying that small drifts could be handled incrementally, but very large changes would require regenerating the model.

In addition to handwriting recognition and document similarity, some in Hinton's group have done quite well using these models in the Netflix contest for movie recommendations (PDF of ICML 2007 paper).

On a lighter note, that is the back of Peter Norvig's head that we see at the bottom of the screen for most of the video. We get two AI gurus for the price of one in this talk.

BellKor ensemble for recommendations

The BellKor team won the progress prize in the Netflix recommender contest. They have published a few papers ([1] [2] [3]) on their ensemble approach that won the prize.

The first of those papers is particularly interesting for the quick feel of the do-whatever-it-takes method they used. Their solution consisted of tuning and "blending 107 individual results .... [using] linear regression."

This work is impressive and BellKor deserves kudos for winning the prize, but I have to say that I feel a little queasy reading this paper. It strikes me that this type of ensemble method is difficult to explain, hard to understand why it works, and likely will be subject to overfitting.

I suspect not only will it be difficult to know how to apply the results of this work to different recommendation problems, but also it even may require redoing most of the tuning effort put in so far by the team if we merely swap in a different sample of the Netflix rating data for our training set. That seems unsatisfying to me.

It probably is unsatisfying to Netflix as well. Participants may be overfitting to the strict letter of this contest. Netflix may find that the winning algorithm actually is quite poor at the task at hand -- recommending movies to Netflix customers -- because it is overoptimized to this particular contest data and the particular success metric of this contest.

In any case, you have to admire the Bellkor team's tenacity. The papers are worth a read, both for seeing what they did and for their thoughts on all the different techniques they tried.

Sunday, December 09, 2007

Facebook Beacon attracts disdain, not dollars

I have been watching the uproar over Facebook Beacon over the last couple weeks with some amusement.

The system was intended to aggregate purchase histories from some online retailers, a poorly thought out attempt to deal with the lack of purchase intent that makes it difficult for Facebook to generate much revenue from advertising.

Om Malik calls ([1] [2] [3]) Facebook Beacon a "privacy nightmare", a "fiasco", and a "major PR disaster". He goes on to write that, even after the latest changes, "I don't think it's easy to trust Facebook to do the right thing."

Dare Obasanjo says "Facebook Beacon is unfixable" because "affiliate sites are pretty much dumping their entire customer database into Facebook ... without their customers permission" and accuses Facebook of "violations of user privacy to make a quick buck."

Eric Eldon at VentureBeat writes that Facebook is dealing with "a revolt against the feature, because it sends messages to your friends about your purchase and other online behavior .... whether or not you are logged in to Facebook and whether or not you have approved any data sharing."

The NYT quotes one Facebook user as saying, "Just because I belong to Facebook, do I now have to be careful about everything else I do on the Internet?" and quotes another as saying, "I feel like my trust in Facebook has been violated."

Facebook faces a hard challenge here. Facebook users are not coming to Facebook thinking of buying things. Because of this lack of commercial intent, most advertisements are likely to be perceived as irrelevant and useless to Facebook users. That will lead to low value to advertisers and low revenues for Facebook.

So, Facebook will struggle desperately to get the revenues promised by their $15B valuation, doing things that almost certainly will annoy and anger their users.

However, Facebook's success largely is due to benefiting from a fad. People flock to whatever social networking site all their friends are using. It wasn't always Facebook. MySpace was once considered by some crowds to be the place to be.

I wonder if all the annoying things that Facebook starts to do with its advertising will be what makes Facebook become uncool. It may be what makes Facebook's fickle audience want to find some new place to hang, somewhere that doesn't suck, and makes Facebook become yesterday's news.

One Wikipedia to rule them all

John Battelle notes a study that reports:
In December 2005 ... 2% of the [top] links proposed by Google and 4% of those proposed by Yahoo came from Wikipedia.

Today 27% of Google's results on the first link alone come from Wikipedia, as do 31% of Yahoo's.
Nick Carr once wrote of this trend, saying:
Could it be that, counter to our expectations, the natural dynamic of the web will lead to less diversity in information sources rather than more?

And could it be that Wikipedia will end up being Google's most formidable competitor? After all, if Google simply points you to Wikipedia, why bother with the middleman?
Update: A week later, the NYT reports that Google started testing a "Wikipedia competitor" called Knol. Google VP Udi Manber "said the goal of Knol was to cover all topics, from science to medicine to history, and for the articles to become 'the first thing someone who searches for this topic for the first time will want to read.'"

Google shared storage and GDrive

The WSJ reports that "a service that would let users store on its computers essentially all of the files they might keep .... could be released [by Google] as early as a few months from now."

Many others have launched products that provide limited storage on the cloud, including AOL, Microsoft, and Yahoo, but Google appears to have an unusual take on it, backing up all of the users files and providing search over them.

One way to view this might be as an extension of the "search across computers" feature of Google Desktop Search. That feature already lets "you to search your home computer from your work computer" by copying an index of the "files that you've been working with recently" and "your web history" to Google, but is limited to only a fraction of your desktop files.

From the sound of the WSJ article -- "a service that would let users store on its computers essentially all of [their] files" -- Google not only will let Google Desktop Search indexes all of your files, but also will copy the original file to the Google cloud.

Please see also my older posts ([1] [2]) on Google's GDrive.

Please see also Philipp Lenssen's post of screen shots from a leaked copy of a version of GDrive (codenamed Platypus) that is only available inside of Google. It apparently replicates and synchronizes all your files across multiple machines.

Friday, December 07, 2007

Google Reader feed recommendations

Googler Steve Goldberg announces that Google Reader has launched feed recommendations based on "what other feeds you subscribe to, as well as your Web History data."

My Google Reader recommendations are quite good. My top recommendations are Seattlest, Slog (from The Stranger, a Seattle free newspaper), Natural Language Processing Blog, Machine Learning etc, and Daniel Lemire's blog.

Pretty hard to complain about those. Nice work, Nitin Shantharam and Olga Stroilova.

It is worth noting that Bloglines has had recommendations for some time, but their recommendations suffer badly from a popularity bias. For example, my top Bloglines recommendations are Slashdot, Dilbert, NYT Technology, Gizmodo, and BBC World News.

[Found via Sriram Krishnan]

Saturday, November 17, 2007

Who cares about grandma?

Jeremy Crane at Compete.com reports that
The top 1% of searchers performs a full 13% of all searches in a given month.

If you extend this to the top 20% the number of queries increase to roughly 70%.
John Battelle argues that the focus on making things easy for common users -- make it work for grandma, as I frequently advocate -- could be misguided.

Since power users make the majority of searches, a search engine that targets power users could attract a majority of searches without attracting the majority of visitors.

However, there is some debate in the comments to John's post about whether Jeremy measured the right thing. What matters most is not the number of searchers, but the ad revenue from those searches.

It is unclear whether these power users who are making the majority of searches are actually the most profitable visitors. Some commenters argue, only anecdotally, that power users may be the least likely to click on ads.

It is an interesting question and one that begs for hard data. Does anyone know if 20% of searchers generate 70% of advertising revenue? Is that 20% is the same 20% that does the 70% of searches? Alternatively, is there is a negative relationship between number of searchers performed by a user and the average revenue per search?

Update: At least for banner advertising, AOL EVP Dave Morgan apparently has some data, writing:
Ninety-nine percent of Web users do not click on ads on a monthly basis. Of the 1% that do, most only click once a month. Less than two tenths of one percent click more often. That tiny percentage makes up the vast majority of banner ad clicks.

Who are these “heavy clickers”? They are predominantly female ... older ... [and] Midwesterners ... They look at sweepstakes far more than any other kind of content. Yes, these are the same people that tend to open direct mail and love to talk to telemarketers.

What does all of this mean? It means that while clickers may be valuable audiences, they are by no means representative of the Web at large.
[Found via Danah Boyd via Jeremy Pickens]

Personalizing the newspaper

I do not agree with everything AP CEO Tom Curley said in his Nov. 1 speech, but I am going to shamelessly pick out the parts I agree with in my excerpts below.

First, some excerpts on personalized news:
The perfect paper or newscast is becoming possible -- at least in the reader's or viewer's eyes. What is it you really want to know? We can personalize content now.

We’re not stuck on those 15-ton behemoths that miraculously manufacture a one-size-fits-all package over several hours that gets delivered over even more hours at great cost or captive of a 22-minute time slot engineered to reach a vast range of content tastes.
The economies of scale with mass production of print newspapers or television broadcasts are much smaller on the Web. On the Web, we have the opportunity to print a different newspaper for each reader, giving each reader a personalized front page.

Next, some excerpts on personalized advertising:
The structure for advertising is changing from mass to targeted.

When you drop a cookie on someone in the digital space, the ads you serve that viewer become up to 200 times more valuable ... The future is about serving ads to people, not to pages or programs.
Offline, we have no opportunity to show different advertisements to different people. The newsprint page, the TV broadcast, the billboard, all are static. Online, we can identify each viewer of the ad space and show something that is likely to be relevant (and maybe even helpful) to that viewer.

Just like Amazon shows a different page to each user -- a store for every customer -- newspapers should build a different page for each reader. Newspapers have gone far too long trying to apply the old static offline model to the online world.

Please see also my Nov 2004 post, "It's the content itself", on a much older speech by Tom Curley calling for personalized news.

[Tom Curley speech found via TechDirt]