Saturday, July 31, 2004

Google's alternative computing platform

Is Google trying to replace the desktop PC? From an article by Rajesh Jain:
    What Google has done is to build an alternative computing platform. This is becoming obvious as it starts our to roll out various services which go beyond just search: a shopping service, social networking, a blogging platform, email with a difference (not to mention plenty of storage), and a local search/yellow pages engine.

    Rick Skrenta [CEO of Topix.net] had this to say: "Google is a company that has built a single very large, custom computer. It's running their own cluster operating system. They make their big computer even bigger and faster each month, while lowering the cost of CPU cycles. It's looking more like a general purpose platform than a cluster optimized for a single application. While competitors are targeting the individual applications Google has deployed, Google is building a massive, general purpose computing platform for web-scale programming."

    Tim O'Reilly [CEO of O'Reilly Media] takes it further: "In a brilliant Copernican stroke, [Google's] gmail turns everything on its head, rejecting the personal computer as the center of the computing universe, instead recognizing that applications revolve around the network as the planets revolve around the Sun. But Google and gmail go even further, making the network itself disappear into the universal virtual computer, the internet as operating system."
I'm not sure. Sounds a lot like the hype surrounding Netscape in the late 1990's. But, if the Windows desktop faces even the remote possibility of being threatened, I can see why Microsoft is responding so aggressively.

Friday, July 30, 2004

ipo.google.com

Interesting tidbits from the transcript on the only recently available IPO site on Google.

Larry Page (Founder) on search:
    You can find a lot of things you're interested in using Google, but we hope to make it much, much better over time, and to really understand the query that you type, to really understand all the information that's available, and while there's a lot of information on the web, not all the information in the world is currently on the web, and so making more information available, understanding it better, understanding what you want better, are all important goals for us ... You'll have an easier time finding the information you're interested in.

Eric Schmidt (CEO) on Adwords:
    Through Google AdWords, advertisers are able to deliver relevant ads cost-effectively to Internet users. The businesses these advertisers are building with Google, as a result of effective targeting, are changing the way advertising works and the way advertisers approach their markets. For example, we don't help advertisers find 24 to 36 year old males. Instead, we help them find consumers interested in purchasing flat panel televisions.

    Unlike the earlier Internet advertising efforts, we didn't just show any ad along with the search. We used special technology invented by Google, to take a search term and figure out which ads were most likely to be relevant. Whereas people tend to ignore untargeted ads, we found that people actually like these ads because they provide additional, relevant information ... That's really the secret of why the model has worked so well for us. We found a way to make advertising useful, not annoying.

George Reyes (CFO) on AdSense:
    AdSense for content focuses on serving targeted and relevant ads based upon the content that the user is reading. Similar to AdSense for search, we get paid each time a user clicks on an ad, and we share the majority of that revenue with our advertising network member. Our AdSense program has become our largest source of revenue, generating about 50% of our revenues through the first half of 2004, up from 24% in 2002.

To summarize, Larry says search is all about relevance (understanding the query and the data to help you find what you need). Eric says advertising is all about relevance (that advertising should be useful and relevant, not obtrusive and annoying). And George confirms my prediction that AdSense (ads placed on other websites) is the key to Google's revenue growth.

So, where does personalization fit into all this? Personalization is all about relevance, recognizing that what is relevant to me isn't the same thing that is relevant to everyone else. If Larry and Eric want relevance, they'll be looking to personalization.

Microsoft's IP strategy

The NYT reports that Microsoft is more aggressively pursuing patents:
    Microsoft said on Thursday that it planned to increase its storehouse of intellectual property by filing 50 percent more patent applications over the next year than in the previous 12 months.
Some are concerned about how these patents might be used. From the NYT article:
    Microsoft, the world's largest software company, increasingly regards the legal protection of its programming ideas as essential to safeguarding its growth opportunities ... Microsoft's stepped-up patent program, analysts say, will be watched closely in the industry to see if the company uses it mainly as a defensive tactic or as an offensive weapon to try to slow the spread of open source products.
And from a CNet article:
    Hewlett-Packard on Tuesday sought to distance itself from a June 2002 memo in which an HP executive said Microsoft planned to use patents as the basis for a legal attack on open-source software.

    "Basically, Microsoft is going to use the legal system to shut down open-source software," said Gary Campbell, then vice president of strategic architecture in HP's office of the chief technology officer, in a memo to several HP executives. "Microsoft could attack open-source software for patent infringements against (computer makers), Linux distributors, and, least likely, open-source developers."

Thursday, July 29, 2004

Microsoft "multisearch"

Ina Fried at CNet reports on Microsoft's blend of desktop and web search:
    Microsoft revealed the progress it has made in building search technology on Thursday when it demonstrated a tool that can comb both the Internet and a PC's hard drive. The technology is designed to quickly look through a hard drive, finding all the matches for a word from within documents, e-mails and even e-mail attachments. The version [Yusuf] Mehdi presented also returned Web results on the right side of the page.
Sounds like this is more than just Lookout.

[Thanks, Gary, for pointing out the article.]

Wednesday, July 28, 2004

Why personalized news?

Why might a personalized news site be more interesting and useful than a manually edited news site?

The problem with a manually edited front page is that everyone sees the same thing. While some top stories about big events are important for everyone to see, picking news stories otherwise is an effort to appease some mishmash of the interests of all readers. It's a compromise that results in mediocrity.

Personalized news focuses you in on the news that is important to you. Important top stories will still appear, but the site also surfaces stories that are important just to you. Have an interest in Linux? Findory News will learn that interest and emphasize news related to Linux. Never interested in sports? Findory News will adapt and deemphasize sports stories.

Personalized news provides a different front page to every reader. It focuses on your interests in a way that is impossible to reproduce by manually editing a page. By uncovering interesting articles from thousands of news sources, it will help you discover news you otherwise would have missed.

Try it out! You might not realize what news you're missing every day.

Tuesday, July 27, 2004

MSN Newsbot review

Microsoft has launched their personalized news product in the US. How does it compare to Findory News?

Using MSN Newsbot provides enough data to do some educated speculation about the system. MSN Newsbot does instantly keep track of articles read and use them immediately to change the small box of "Personalized News" headlines in the upper right of the front page. Aside from the small box of personalized headlines, the rest of the page appears to be unpersonalized. Some articles I read did not appear to be recorded. For example, I did a search on "newsbot" and clicked on three articles, only to find that none of the articles were recorded in my history or used for personalization. I also managed to lose my entire history once for no reason I could discover.

It's difficult to determine the underlying algorithms from inspecting the behavior of the site, but there appears to be strong evidence that it is mostly based on subject categories. Reading an article on business in Korea caused top headlines from Asia to appear. Surprisingly, even after deleting the article from my history, Asia top stories continued to be selected. Clicking on the "Why?" link gave the explanation that the personalized stories were picked because they are from the category "Asia-Pacific Latest".

Similarly, reading an article on Google's IPO caused the personalization to show me more top headlines from "Business:General" and "Business:Financial". Reading a science article on the effects of caffeine just produced more general science articles.

Using subject-based profiles is a well-known method of doing personalization, but it also has well known problems. In particular, the personalization is not specific -- for example, showing just general business headlines -- and tends to pigeonhole people -- showing a reader only business stories and not picking up other cross category interests. While it does have the advantage of being simple, experience in my past life shows that the predictive accuracy of this method is an order of magnitude lower than more fine-grained personalization techniques.

If it is true that MSN Newsbot is merely using subject classifications for its personalization, Findory's personalization technology is considerably more advanced. Findory's algorithms combine statistical analysis of the article text and of users who viewed the articles with information about articles you previously viewed. Our personalized news is fine-grained. Our personalization is targeted closely to your interests while maintaining enough serendipity to enhance discovery. We help you read the news more efficiently and find articles you otherwise would miss. There is still nothing else like it out there.

Monday, July 26, 2004

MSN Newsbot launched

Chris Sherman reports that Microsoft's personalized news site, MSN Newsbot, has launched in the US. Interesting to see how it compares to Findory News.

Update: The press release from Microsoft.

Update: Press coverage of the launch seems to be focusing on MSN Newsbot as a "Google News killer". Certainly true that this represents a new front in the search war between Google, MSN, and Yahoo.

Update: Some good coverage on Poynter E-Media Tidbits and ResourceShelf.

Re-recruiting your knowledge nomads

Niall Kennedy comments on an excellent HBS article, "High Turnover: Should You Care?". Some excerpts from the article on commitment and productivity:
    Length of time in a company is the most common way of measuring employee commitment. And it is the least interesting and least helpful approach for managers. Far more important is the quality and quantity of work someone does when in a company.
And on retention and job satisfaction:
    Employers do often look at turnover as only a bad thing, [but] research has demonstrated that some turnover is healthy, indeed essential to organizational well being. But even more important is what managers do when they see turnover as a bad thing. To drive it down they often employ the wrong remedies.

    Managers who were asked to identify ways to retain workers came back with action steps like "increase salary" and "change his or her title." These are small changes [that] may keep an employee in a company for a couple of months, but they will not hold an employee for long, and little productivity will be gained. The managers we asked to identify ways to elicit commitment proposed deeper and more individualized action steps, like "find out what challenges make him or her tick" and "provide opportunities for learning on the job." ... Most employees seek to be valued and engaged.
Turnover is a symptom, not the problem. Keep people challenged, learning, and happy. It's good advice.

A lonely and difficult road

Warren Schultz at CareerJournal has a good reality check for aspiring entrepreneurs in his latest column, Entrepreneurship Is Often A Lonely and Difficult Road. While I think he underestimates the benefits of starting your own business -- my own experience has been fantastic -- Warren is right to warn of the costs and risks.

Sunday, July 25, 2004

Technorati's priorities

Steve Rubel criticizes Technorati for emphasizing marketing over engineering:
    I love Technorati, but this smells like dot com spirit all over again. Where's the moolah coming from to support a PR team of five? Hiring a PR firm before you can handle demand and squash bugs is looking for trouble. Hope they are ready for all the added attention.
I've personally had some problems with the performance, stability, and completeness of Technorati's search. I'm hopeful that Adam Hertz, Technorati's new VP of Engineering, will be addressing these issues quickly.

Update: After being frustrated by slow and failed Technorati searches again today, I think I'm starting to agree with Jason Calacanis.

Update: David Sifry (CEO of Technorati) acknowledges and apologizes for their scaling problems and outages.

Saturday, July 24, 2004

BugMeNot "registration form"

BugMeNot, a site I've mentioned ([1] [2]) that allows people to bypass mandatory registration requirements on websites, has an amusing joke registration form. As Cory Doctorow says, it cleverly "exemplifies many of the critical problems with registration on the Web."

Friday, July 23, 2004

Findory mentioned on Poynter

Steve Outing writes about Findory News:
    Personalizing news websites is a nice idea, but existing attempts I've found to be less than the ideal. Most require user registration, and then you can specify which sections (Business, Sports, etc.) you want on your home page. But there's another way, as demonstrated by Findory.com, a news site that debuted in March. It's a news aggregator (like Google News), and it's personalized (like My Yahoo!). The cool part about Findory is that whenever you click on an article link (which brings up the story from a news site, just like any other news portal) the site remembers what you've read and decides what other stories you might be interested in based on your previous clicks. It learns your interests throughout your current browsing session and subsequent ones, and presents headlines based on your past clicking behavior. (For example, I clicked on a Tour de France article; when I returned to the homepage of Findory.com, the article ranking had immediately included more bicycling stories out front.)
Steve Outing is a senior editor at the Poynter Institute for Media Studies, and an interactive media columnist for Editor and Publisher, a journal on the newspaper industry.

Thursday, July 22, 2004

Slashdot moderation and reputation

I've been frustrated with Slashdot moderation lately. The current system doesn't properly emphasize the most informative and interesting comments.

While the suggestions in Slash(dot) and Burn may help, I'm convinced that more is needed. In particular, I think the comment and moderation system needs to do more with reputation.

Currently, Slashdot has a simple reputation system called karma. Users with karma over a threshold have a higher initial score on their comments. High karma users also are occasionally given a few moderation points to raise or lower the score of other comments by +-1. Comments have a score in the range [-1, 5] and users can elect to filter all comments below a threshold.

But why these thresholds? Why not have everything work as a function of karma? For example, the starting score of a comment from user with low karma could be 1.15, from a medium karma user 1.31, and very high karma 2.38. Allow almost all users with positive karma could do at least some moderating, but moderations from a low karma user should barely nudge the score, perhaps by as little as +- 0.04, while a moderation from a very high karma user should move the score by +- 1.33.

Since high karma users mostly get their karma from posting interesting comments, giving high karma users more influence should improve the quality of discussions. In addition, using the full range [-1,5] instead of only integer values will allow more subtle differentiation between comments.

What do you think? Would this work? Are there other ways Slashdot moderation could be improved?

Wednesday, July 21, 2004

Growth and the future at Yahoo Search

Some tidbits on Search Engine Watch about Yahoo Search:
    With the release of its new search engine, Yahoo now powers over half of the US web searches - this is a dramatic shift in the market share within the industry. Yahoo now has 260 million users world wide and 100 million registered users.

    Yahoo sees personalized search as the future focus. The goal of personalization is to better understand the user intent. Currently, people have to type in extra words in their query to be more specific to get the results they want. With Personalized search, the search engine delivers relevant results with fewer words. For example, if people want a haircut and Yahoo knows that they live in midtown New York, Yahoo would be able to automatically supply haircutters in that area.
Yahoo's large registered user base is an advantage for personalization. Signed-in customers will not lose all their data if they lose their cookie. Signed-in users can see the same personalization from multiple computers. And many users may have provided demographic data and information about their interests.

The new command line

"The Location Field is the New Command Line" is a cleverly titled essay on web applications (like Google's GMail) and the disruptive impact they could have on Microsoft Windows.

Tuesday, July 20, 2004

Registration required

Rachel Metz at Wired writes about the annoyance of registration requirements at news web sites and tools like BugMeNot that thwart them. The article contains interesting quotes from people in the industry on why they require registration:
    Elaine Zinngrabe, general manager of latimesinteractive, which runs the Los Angeles Times' website, said the newspaper began requiring online user registration in June 2002 as a way to learn more about its readers and, it hopes, to drum up more advertising on the site. The Times asks readers to reveal things like their ZIP code, age, gender and income.

    Dipik Rai, a business manager with Knight Ridder Digital who runs online registration for some of the company's newspapers ... said Knight Ridder Digital is very upfront about why it's gathering data, which it uses to figure out who's using its site and to target advertising.
Both Elaine and Dipik claim that the reason they need registration is to understand their audience and for online advertising. There are three issues with this:
  1. Random sampling: Acquiring a demographic profile of your audience only requires a random sample of your customers, not forcing every single reader to fill out a lengthy form.
  2. Effective targeted advertising: I suspect a web site can generate higher targeted online advertising clickthroughs and revenue by matching ads to user behavior (the articles the reader has read) than from noisy and coarse-grained data on age, zip code, and income.
  3. Cost of registration: Required registration repels some visitors from your website, so you lose traffic and advertising dollars. The cost of lost revenue needs to be part of the cost-benefit analysis of required registration.
Read more:

Monday, July 19, 2004

They think they can

An interesting article on innovation in smaller search engines. Mentions of Topix.net, Find.com, Vivisimo, and Eurekster, plus a few interesting comments from John Battelle.

Sadly, no mention of Findory News or Findory Blogory.

Netflix

Two interesting articles on Netflix this morning:

Phillip Torrone on Engadget argues that Netflix should provide open APIs to its system like Google and Amazon do.

Hacking Netflix talks about Blockbuster's leaked beta test of a Netflix imitator.

My biggest issue with Netflix is that their recommendations don't seem very effective. They're obviously trying to bias the recommendations heavily to the less frequently rented items in the back catalog, but they bias so heavily that the recommendations are never interesting to me.

It's a real shame. If the recommendations were any good, I might be reluctant to switch to a competitor, since I'd lose my nifty recommendations by making the switch. As it is, I don't have much reason to stick with Netflix, especially as they raise prices and appear to be having increasing problems with availability.

Saturday, July 17, 2004

Forrester Research picks Yahoo

Forrester Research predicts that use of Yahoo Search will exceed Google in Q1 2005. George Colony, Forrester chair and CEO points out that:
    Great search site has three components: personalization, presentation, and quality of service. Of all the search engines out there, Yahoo is the only player that gets all three.
I'm not sure I agree that Yahoo is as good as Google on presentation or quality of service. But, Google's attempt so far at personalization has been weak. Yahoo clearly considers ([1] [2]) personalization to be the way for them to compete.

Friday, July 16, 2004

InfoWorld on RSS growing pains

Chad Dickerson at InfoWorld talks about scalability issues with RSS:
    InfoWorld.com sees a massive surge of RSS newsreader activity at the top of every hour, presumably because most people configure their newsreaders to wake up at that time to pull their feeds. If I didn't know how RSS worked, I would think we were being slammed by a bunch of zombies sitting on compromised home PCs. Our hourly RSS surge has all the characteristics of a distributed DoS attack, and although the requests are legitimate and small, the sheer number of requests in that short time period creates some aggravating scaling issues.
I commented earlier on a Wired article that expressed similar concerns about the scalability of the polling architecture of RSS.

There's a variety of ways to deal with this issue. The solution Chad seems to be suggesting is to randomize request times so that there aren't big spikes in traffic every hour at the hour. That's certainly a good idea. Clients should also respect the ttl (polling at the interval that is listed in the feed), support conditional GET, and handle 304 (not modified) responses to minimize the number of requests they make for the full feed.

But the primary solution will end up being caching. With the exception of personalized RSS feeds, RSS feeds easily can be cached. Web-based RSS readers like Bloglines and My Yahoo already only read the RSS feed once, cache it, and display it to multiple readers. But popular RSS feeds are also easily proxy cached just like web pages, reducing the load on the original source servers.

Thanks, RSS Weblog, for pointing out the InfoWorld article.

Update: An interesting discussion of RSS scaling over on Jeremy Zawodny's blog.

Update: Chad Dickerson follows up with an article summarizing all the suggestions he received. It's mostly the same as my suggestions above, but still probably worth skimming.