The metasearch engine Dogpile released a snazzy little Flash app they call "Missing Pieces" that lets you query Google, Ask Jeeves, and Yahoo simultaneously and see the overlap between results. Well designed, fun little app.
Chris Sherman has a detailed review of Dogpile's recent redesign and new tools.
Friday, May 13, 2005
Thursday, May 12, 2005
Millions of feeds
Alex has the scoop on Findory's millions of feeds, all available by RSS or displayed in our nifty inline widget.
Get a feed for Tech news, a personalized selection of articles from political blogs, even all news or weblog articles that match specific keywords like "Microsoft", "Seattle Sonics", "San Diego", or "Star Wars". There are millions of possible combinations, a feed for every taste.
And you can put news and blog articles on any topic you want right on to your website. Want to show the latest news about Linux on your blog? Go to our inline page, type "Linux" into the search box, and copy the Javascript code it gives you into your blog page.
Just like our RSS feeds, inline works for millions of possible categories and keyword searches. You can even show your weblog readers a view of your personalized Findory page by putting up your personalized version of inline. If you need to match the look and feel of your website, just click "Show optional style code" and customize inline as much as you want.
There are three examples -- related blog posts to Geeking with Greg, a snippet of the news from my personalized front page, and news for "Google" -- in the right side column on my weblog. They are customized to match my weblog style. They update as new news comes in. The personalized version ("My Findory News") even changes in real time as I read new articles on Findory. Very cool.
Get a feed for Tech news, a personalized selection of articles from political blogs, even all news or weblog articles that match specific keywords like "Microsoft", "Seattle Sonics", "San Diego", or "Star Wars". There are millions of possible combinations, a feed for every taste.
And you can put news and blog articles on any topic you want right on to your website. Want to show the latest news about Linux on your blog? Go to our inline page, type "Linux" into the search box, and copy the Javascript code it gives you into your blog page.
Just like our RSS feeds, inline works for millions of possible categories and keyword searches. You can even show your weblog readers a view of your personalized Findory page by putting up your personalized version of inline. If you need to match the look and feel of your website, just click "Show optional style code" and customize inline as much as you want.
There are three examples -- related blog posts to Geeking with Greg, a snippet of the news from my personalized front page, and news for "Google" -- in the right side column on my weblog. They are customized to match my weblog style. They update as new news comes in. The personalized version ("My Findory News") even changes in real time as I read new articles on Findory. Very cool.
Wednesday, May 11, 2005
Googleball
Google acquires Dodgeball, a two-person mobile social networking startup.
More details from Michael Bazeley and Gary Price.
Update: About two years later, both the founder and the second employee of Dodgeball leave Google, saying:
More details from Michael Bazeley and Gary Price.
Update: About two years later, both the founder and the second employee of Dodgeball leave Google, saying:
It's no real secret that Google wasn't supporting dodgeball the way we expected. The whole experience was incredibly frustrating for us.Perhaps Google is not as proficient at handling these small acquisitions as they claim to be.
Tuesday, May 10, 2005
Answers.com about Shopping.com
Gary Price posts about Answers.com's deal to integrate product information from Shopping.com.
Answers.com is a metasearch engine that hits specialized databases to try to directly answer your query rather than returning a list of web pages that might contain the answer to your query. More information in my earlier post, "Answers.com launches".
Google.com recently switched their word definitions links from Dictionary.com to Answers.com. I'm curious if this new deal could put some new strain on that relationship. After all, Shopping.com and Google's Froogle are similar metashopping search engines.
As the overlap increases, helping GuruNet's Answers.com may start to be less and less in Google's interests. Google's word definitions could just as easily be provided by Wikipedia directly or by a homegrown combination of dictionary and Wikipedia content.
Answers.com is a metasearch engine that hits specialized databases to try to directly answer your query rather than returning a list of web pages that might contain the answer to your query. More information in my earlier post, "Answers.com launches".
Google.com recently switched their word definitions links from Dictionary.com to Answers.com. I'm curious if this new deal could put some new strain on that relationship. After all, Shopping.com and Google's Froogle are similar metashopping search engines.
As the overlap increases, helping GuruNet's Answers.com may start to be less and less in Google's interests. Google's word definitions could just as easily be provided by Wikipedia directly or by a homegrown combination of dictionary and Wikipedia content.
Saturday, May 07, 2005
Profiting from the long tail
In "Profiting from Obscurity", The Economist talks about how companies profit from helping customers find products in the long tail. Some excerpts:
- [The long tail] is a shift from mass markets to niche markets, as electronic commerce aggregates and makes profitable what were previously unprofitable transactions.
How can people find content they want when it is buried far down the tail? Already, a number of mechanisms have emerged, based around user recommendations. Perhaps the best known is "collaborative filtering", in which purchase histories are analysed to work out what else is likely to interest the buyer of a particular product ("Customers who bought this item also bought...", as Amazon puts it). This approach allows users to navigate from hits that they know they like to more obscure titles further down the tail.
Many successful online businesses, such as Amazon, Rhapsody or the iTunes Music Store, already exploit the effects of the long tail. So too do other internet companies, such as Google (which makes money not just by selling adverts to big firms, but also by placing obscure adverts alongside obscure web pages) and eBay (which aggregates low levels of demand for obscure products to make a huge business).
Friday, May 06, 2005
Controversy about Google Web Accelerator
Many have been talking about Google Labs' latest release, Google Web Accelerator. It promises to reduce wait times when web browsing mostly by a combination of caching and pre-fetching.
Unfortunately, it appears to have some issues. First, as Nathan Weinberg writes, it appears that Google Web Accelerator is sometimes serving up the wrong cached page, showing you a page with someone other than you logged into site, for example.
Second, as Fons Tuinstra reports, Google Web Accelerator effectively acts like a proxy and, among other things, allows users to bypass China's firewall. While I'm no supporter of that firewall, I suspect China will not be happy about this development. The response is unlikely to be favorable to Google.
Finally, many have pointed out the privacy implications of Google knowing about most pages each user has ever visited anywhere on the web and the full contents of those pages. Google may have earned a lot of trust, but this is a big step, and one that is likely to cause some concern.
But keep the big picture in mind here. Not only is Google providing a free product that can save people time if they choose to use it, but also properly anonymized information about what pages people are visiting could be used by something like TrustRank to help reduce web spam and improve the relevance rank of search results. In the end, Google is helping people find the information they need more quickly and efficiently.
Update: Hmm... I have to say I'm feeling a lot less charitable toward Google Web Accelerator after seeing Jason Fried's post about how Google Web Accelerator can delete data and have other undesirable behaviors. It looks like a lot of sites, including Findory, are going to have to go through the effort of explicitly disabling Google's prefetch. Ugh, what a mess.
Unfortunately, it appears to have some issues. First, as Nathan Weinberg writes, it appears that Google Web Accelerator is sometimes serving up the wrong cached page, showing you a page with someone other than you logged into site, for example.
Second, as Fons Tuinstra reports, Google Web Accelerator effectively acts like a proxy and, among other things, allows users to bypass China's firewall. While I'm no supporter of that firewall, I suspect China will not be happy about this development. The response is unlikely to be favorable to Google.
Finally, many have pointed out the privacy implications of Google knowing about most pages each user has ever visited anywhere on the web and the full contents of those pages. Google may have earned a lot of trust, but this is a big step, and one that is likely to cause some concern.
But keep the big picture in mind here. Not only is Google providing a free product that can save people time if they choose to use it, but also properly anonymized information about what pages people are visiting could be used by something like TrustRank to help reduce web spam and improve the relevance rank of search results. In the end, Google is helping people find the information they need more quickly and efficiently.
Update: Hmm... I have to say I'm feeling a lot less charitable toward Google Web Accelerator after seeing Jason Fried's post about how Google Web Accelerator can delete data and have other undesirable behaviors. It looks like a lot of sites, including Findory, are going to have to go through the effort of explicitly disabling Google's prefetch. Ugh, what a mess.
Tuesday, May 03, 2005
Seattle Times on small search companies
Findory was mentioned in a Seattle Times article covering search companies in the Seattle area. There's a brief discussion of the rapid growth of our tiny personalization startup. And an amusing picture of me that I think truly captures the inner geek.
UW CS Professor Oren Etzioni has a great quote in the article: "Could we build a search engine that learns from the Web? ... Imagine a program that is running and constantly learning from the Web and getting smarter over time."
It's an excellent vision, the dream of many AI researchers, understanding the vastness of knowledge stored in the Web.
UW CS Professor Oren Etzioni has a great quote in the article: "Could we build a search engine that learns from the Web? ... Imagine a program that is running and constantly learning from the Web and getting smarter over time."
It's an excellent vision, the dream of many AI researchers, understanding the vastness of knowledge stored in the Web.
Monday, May 02, 2005
Web spam and TrustRank
I finally managed to get a good look at the "Combating Web Spam with TrustRank" paper by Gyongyi et al this weekend.
TrustRank takes a manually designated set of good or bad pages and propagates that information across the link graph. It's an interesting modification to PageRank. Definitely worth a read.
The paper describes a manual process for determining the seed set of trusted sites. I'm curious what we'd find by instead analyzing user behavior. For example, we could consider websites used over the past month by trusted people to be trusted. That is, trusted sites would be the sites the community uses and trusts.
Noisier data, to be sure, but there's a sea of data here, enough that we should be able to be robust to the noise and discover the wisdom hidden within. Ah... So many interesting possibilities with this kind of juicy data.
By the way, if you're interested in web spam, don't miss "Web Spam Taxonomy", also by Zoltan Gyongyi and Hector Garcia-Molina. It's a light paper that describes many of the devious techniques used by web spammers.
Update: There appears to be a recent March 2006 technical report on this TrustRank work, "Link Spam Detection Based on Mass Estimation" (PDF).
TrustRank takes a manually designated set of good or bad pages and propagates that information across the link graph. It's an interesting modification to PageRank. Definitely worth a read.
The paper describes a manual process for determining the seed set of trusted sites. I'm curious what we'd find by instead analyzing user behavior. For example, we could consider websites used over the past month by trusted people to be trusted. That is, trusted sites would be the sites the community uses and trusts.
Noisier data, to be sure, but there's a sea of data here, enough that we should be able to be robust to the noise and discover the wisdom hidden within. Ah... So many interesting possibilities with this kind of juicy data.
By the way, if you're interested in web spam, don't miss "Web Spam Taxonomy", also by Zoltan Gyongyi and Hector Garcia-Molina. It's a light paper that describes many of the devious techniques used by web spammers.
Update: There appears to be a recent March 2006 technical report on this TrustRank work, "Link Spam Detection Based on Mass Estimation" (PDF).
Seattle Times on the search war
In "Microsoft Learns to Crawl", Kim Peterson at the Seattle Times gives an inside view of MSN Search's efforts in the search war. Some excerpts:
See also Fred Vogelstein's excellent Fortune article on MSN and the search war.
- The search battle is bigger than beating Google. Microsoft is carving its path in the next generation of computing -- one in which search becomes a platform, not a feature ....
Some ... are eyeing perhaps the biggest weapon in Microsoft's arsenal: the operating system. The company is building search into its upcoming operating system, code-named Longhorn and expected next year, and likely will give it a prominent spot in front of the user.
See also Fred Vogelstein's excellent Fortune article on MSN and the search war.
Wednesday, April 27, 2005
Adam Bosworth on simple web services
Daniel Steinberg reports on Adam Bosworth's talk at the MySQL Users Conference. Adam criticized past efforts on web services as being too complicated and advocated simpler techniques:
[via Dare Obasanjo]
- Imagine if you can query any data that is available anywhere in the world ... What this requires is a single, simple, open wire format for items. The format needs to be simple for any P programmer to deliver and any JavaScript programmer to consume ... "Complex things tend to break and simple things tend to work."
RSS 2.0 and Atom will be the lingua franca that will be used to consume all data from everywhere. These are simple formats that are sloppily extensible. Anyone who wants to can use these formats to consume content or to author content.
[via Dare Obasanjo]
Tuesday, April 26, 2005
Yahoo My Web and searching web history
Yahoo launched My Web, an extension to the Yahoo Toolbar that lets you save on Yahoo's servers a copy of any web page you've seen. The idea is that it lets you find web pages you saw once before quickly and easily.
This isn't quite like Seruku, which saves a copy of every web page you've seen on your own computer. Or quite like Google Desktop Search, which searches over every web page you've seen recently (using the browser cache on your computer). Or quite like Filangy, which (sort of) stores every page you've seen on their servers.
Because you have to explicitly push a button to save the page, I'm not sure how many people will use Yahoo My Web. Seems likely to me that, when I'm browsing a particular page, I don't really know if I'll want to find it again. It'll be too late when I decide I really should have pushed that little button.
Cute idea though. Fun to see all the innovation lately.
See also Chris Sherman's review of Yahoo My Web.
Update: Tony Gentile posts a detailed review. He also points out that Yahoo My Web probably should be considered more of a bookmarking tool than a step toward Memex.
This isn't quite like Seruku, which saves a copy of every web page you've seen on your own computer. Or quite like Google Desktop Search, which searches over every web page you've seen recently (using the browser cache on your computer). Or quite like Filangy, which (sort of) stores every page you've seen on their servers.
Because you have to explicitly push a button to save the page, I'm not sure how many people will use Yahoo My Web. Seems likely to me that, when I'm browsing a particular page, I don't really know if I'll want to find it again. It'll be too late when I decide I really should have pushed that little button.
Cute idea though. Fun to see all the innovation lately.
See also Chris Sherman's review of Yahoo My Web.
Update: Tony Gentile posts a detailed review. He also points out that Yahoo My Web probably should be considered more of a bookmarking tool than a step toward Memex.
Press release on Findory growth
Findory just issued a press release on our exponential growth. An excerpt:
- Since the site launched in early 2004, Findory's traffic has more than doubled every three months.
With over one million page views a month, tens of thousands of people rely on Findory's personalized front page to keep up with everything from local politics to world cup rugby.
Over one million articles have been read through Findory.com. Findory's traffic has grown nearly thirty times in the last twelve months.
Monday, April 25, 2005
Improving article pages
Steve Outing at Editor & Publisher writes about improving online news article pages:
[via Simon Waldman]
- More and more people bypass news Web sites' home and section pages .... The best approach is to create an article-page template that serves as a sort of secondary home page ... Give them enough choices to guide them to other important content elsewhere on the site.
According to its Web site's editor, Angus Frame, 41% of globeandmail.com visits now begin on non-hub pages ... "Readers who went straight to a story page had no idea how much was available on globeandmail.com," [Frame said]. "They only saw one story and then had no reason to stick around."
Globeandmail.com's managers decided last year to "improve the story-page experience." [Frame said], "We added valuable, informative links to the right-hand side of the story in a fairly wide column. We turned every story into a mini hub .... Literally overnight daily page views increased by more than 25%, from about 2.3 million page views a day to 3.0 million page views a day."
[via Simon Waldman]
Sunday, April 24, 2005
Findory redesign
Findory launched a redesign of our site today. We affectionately have been calling it the "fuel" release in honor of the caffeine-charged coding frenzies that took place at Fuel Coffee here in Seattle.
Alex did tremendous work combining suggestions from our readers with new features such as popular sources, creating a clean and attractive new look. Our beta testers have been gushing over it. We hope you like it too!
Alex did tremendous work combining suggestions from our readers with new features such as popular sources, creating a clean and attractive new look. Our beta testers have been gushing over it. We hope you like it too!
Friday, April 22, 2005
Trying to Wiki the news
Joanna Glasner at Wired writes about some problems at Wikinews:
See also my earlier post, "Slash(dot) and Burn".
- Operators of Wikinews are finding their mission rife with frustrations and challenges.
The site, an offshoot of Wikipedia, the volunteer-maintained online encyclopedia, is facing pressures its parent organization rarely had to contend with, such as ferreting out fake posts, incorporating original sources and updating coverage to reflect rapidly changing current events.
"In Wikipedia, the writing style of an encyclopedia is more timeless. You can get it right eventually. It's going to be the same article for many years," said Jimmy Wales, Wikipedia's founder. "With a news story, the actual story has a limited lifespan. If it's not neutral, you've got to fix it quickly."
The Wikinews site follows essentially the same set of rules as the Wikipedia encyclopedia, which allows anyone to create entries or edit and correct other people's work.
See also my earlier post, "Slash(dot) and Burn".
Thursday, April 21, 2005
Text statistics on Amazon.com
Nathan Torkington at the new O'Reilly Radar blog points to a new feature on Amazon that can show you the most frequently used words from and various statistics on many books.
For example, here's the concordance and text statistics for Applied Cryptography.
This is cute, but not very useful. I'm not sure I care that Applied Cryptography averages 1.7 syllables per word or that the most common word in the book is "key".
The "books on related topics" on the same page might be more useful. Amazon says the relationships are determined using their new SIPs data. It would be fun to experiment with using this text analysis data to try to find improvements to the accuracy of Amazon's personalization and recommendations.
For example, here's the concordance and text statistics for Applied Cryptography.
This is cute, but not very useful. I'm not sure I care that Applied Cryptography averages 1.7 syllables per word or that the most common word in the book is "key".
The "books on related topics" on the same page might be more useful. Amazon says the relationships are determined using their new SIPs data. It would be fun to experiment with using this text analysis data to try to find improvements to the accuracy of Amazon's personalization and recommendations.
Filangy and searching your web history
Gary Price posts on Filangy, a toolbar that keeps track of every web page you've visited and lets you search over your web browsing history.
One curious feature of Filangy is that your history is stored on their servers, not on your PC. The advantage of this is that your history can be combined and accessed from multiple computers. The disadvantage is that the snapshot of the page Filangy indexes may not be the page you actually viewed.
For example, if you go to Amazon's home page, you'll see a personalized page with content picked just for you. When Filangy retrieves the Amazon page from their servers and stores it, it appears that the page they retrieve will be the generic page, not the page you just viewed.
It's a little strange to have a search of your browsing history that searches pages you never actually saw.
About a year ago, I posted about a similar toolbar called Seruku and the Microsoft Research "Stuff I've Seen" project, both of which try to make it easy to find anything you've found before.
Since then, Google launched its desktop search. Unlike the desktop search tools from MSN and Yahoo, Google Desktop Search indexes your web history (at least the cached last few days of it), a very useful feature.
Search over your browsing history brings us closer toward Memex, the memory extender. In the near feature, you will be able to easily recall anything you've seen on your computer. Filangy, Seruku, MSR's "Stuff I've Seen" and Google Desktop Search are bringing us closer to that future.
One curious feature of Filangy is that your history is stored on their servers, not on your PC. The advantage of this is that your history can be combined and accessed from multiple computers. The disadvantage is that the snapshot of the page Filangy indexes may not be the page you actually viewed.
For example, if you go to Amazon's home page, you'll see a personalized page with content picked just for you. When Filangy retrieves the Amazon page from their servers and stores it, it appears that the page they retrieve will be the generic page, not the page you just viewed.
It's a little strange to have a search of your browsing history that searches pages you never actually saw.
About a year ago, I posted about a similar toolbar called Seruku and the Microsoft Research "Stuff I've Seen" project, both of which try to make it easy to find anything you've found before.
Since then, Google launched its desktop search. Unlike the desktop search tools from MSN and Yahoo, Google Desktop Search indexes your web history (at least the cached last few days of it), a very useful feature.
Search over your browsing history brings us closer toward Memex, the memory extender. In the near feature, you will be able to easily recall anything you've seen on your computer. Filangy, Seruku, MSR's "Stuff I've Seen" and Google Desktop Search are bringing us closer to that future.
Wednesday, April 20, 2005
Google launches search history
Google just launched My Search History. It keeps track of your previous searches and search result clickthroughs, making it easy to find things again that you found in the past.
A9, My Yahoo Search, My Ask Jeeves, and Findory have had this feature for a while. All but Findory require users to sign in to activate the feature. Like Findory and A9 but unlike Yahoo and Ask, Google's search history feature is integrated into the main search on the site.
Keeping search and clickthrough history is a first step toward personalized search. The next big step is to use this data to reorder search results, making the results more relevant to your particular interests and needs.
Personalized search is the future. In his article on Google's new search history feature, Chris Sherman says:
[See also Stefanie Olsen's article at CNet]
Update: Charlene Li says personalized search results are the "Holy Grail of search" and quotes Google Director Marissa Mayer for an example of how Google might personalize search results.
Update: Danny Sullivan posts an interesting comparison chart of search history features from A9, Ask, Eurekster, Findory, Furl, Google, and Yahoo.
A9, My Yahoo Search, My Ask Jeeves, and Findory have had this feature for a while. All but Findory require users to sign in to activate the feature. Like Findory and A9 but unlike Yahoo and Ask, Google's search history feature is integrated into the main search on the site.
Keeping search and clickthrough history is a first step toward personalized search. The next big step is to use this data to reorder search results, making the results more relevant to your particular interests and needs.
Personalized search is the future. In his article on Google's new search history feature, Chris Sherman says:
- Don't expect Yahoo, Ask Jeeves, MSN or AOL Search to stand still. Personalized search has long been touted as one of the holy grails for the industry ... Beginning today with Google's launch of My Search History, I expect to see major leaps ahead in the arena of personalized search -- and that's a good thing.
[See also Stefanie Olsen's article at CNet]
Update: Charlene Li says personalized search results are the "Holy Grail of search" and quotes Google Director Marissa Mayer for an example of how Google might personalize search results.
Update: Danny Sullivan posts an interesting comparison chart of search history features from A9, Ask, Eurekster, Findory, Furl, Google, and Yahoo.
Tuesday, April 19, 2005
Fortune on the search war
Fred Vogelstein at Fortune wrote a great article on MSN and the search war with Yahoo and Google. The full article is subscription only, but here are some selected excerpts.
Microsoft has been having difficulties in the search war:
Microsoft has been having difficulties in the search war:
- Every month it seems as if Google hires away one of Microsoft's top developers ... As of March, roughly 100 Microsofties had left for its search nemesis ... The Google migration has gotten so bad, says a former Microsoft employee, that when he told his bosses and colleagues he was leaving earlier this year, "the first question out of their mouths was 'You're not going to Google, are you?'"
Trying to build a Google killer ... has turned out to be truly humbling for Microsoft. The effort has taken longer, cost more money, and exposed more big-company problems at Microsoft than anyone imagined.
- One reason Google has been rolling out so many new or improved products is that [Google CEO Eric] Schmidt understands that innovation is the only sure edge Google has. The moment Google allows itself to slow, Microsoft could overwhelm it.
- "We need to take search way beyond how people think of it today and just have it be naturally available, based on the task they want to do," [said Bill Gates]. For example, if you wanted to look up a factoid while you were writing a document, you might search for it without ever leaving Word.
All three big search engines are scrambling to find ways to make search more personalized. The thinking is that the more a search engine knows about who is searching, the more accurate the results will be.
Google Maps launches in UK
Just when I get back from London, Google launches Google Maps and Google Local in the UK.
Looks cleaner than the alternative, Streetmap.co.uk. It's great to see more options for navigating the confusing twists and turns of London. I wish I had had it last week!
Looks cleaner than the alternative, Streetmap.co.uk. It's great to see more options for navigating the confusing twists and turns of London. I wish I had had it last week!
Subscribe to:
Posts (Atom)