If a picture is worth a thousand words, how valuable is the ability to find the perfect image of an object from the entire Web? According to a paper delivered by two Google researchers at the International World Wide Web Conference in Beijing last weekend, the search-engine giant may be one step closer to answering that question.
Information scientists Shumeet Baluja and Yushi Jing announced the development of an algorithm, called VisualRank, that generates significantly more relevant image-search results than current results using text-based clues (captions and other words associated with each image).
The goal ultimately is to train computers to move beyond text into the effective identification of “rich content” — the shapes, colors and context of images that humans recognize with little effort.
VisualRank, Baluja and Jing reported, is designed to incorporate ongoing advances in computer recognition into Web search technology. The complicated process blends image-recognition advances with Google’s sophisticated tools for assigning rank and weight to search results.
The net effect, they said, is that within a relatively narrow universe of search results, the algorithm was able to reduce the number of irrelevant results by more than 80 percent.
But as Baluja and Jing freely concede, it is highly impractical to try to identify comparable images among the billions currently stored on the Web. To test its system, the Google team created data sets of images of the 2,000 products most commonly searched for on Google. Team members then assigned a relevance score to images produced by Google’s normal image-search tool and VisualRank.
One of the questions is whether VisualRank has practical market possibilities or is merely a challenging intellectual exercise. As industry observers have pointed out, the Web site Like.com also offers surfers the ability to locate images of similar products by searching for a particular element in each image….