The desire for greater control over how search engines index and display Web sites is driving an effort by leading news organizations and other publishers to revise a 13-year-old technology for restricting access.
Currently, Google Inc., Yahoo Inc. and other top search companies voluntarily respect a Web site’s wishes as declared in a text file known as “robots.txt,” which a search engine’s indexing software, called a crawler, knows to look for on a site.
The formal rules allow a site to block indexing of individual Web pages, specific directories or the entire site, though some search engines have added their own commands.
The new proposal, to be unveiled Thursday by a consortium of publishers at the global headquarters of The Associated Press, seeks to have those extra commands — and more — apply across the board. Sites, for instance, could try to limit how long search engines may retain copies in their indexes, or tell the crawler not to follow any of the links that appear within a Web page.
The current system doesn’t give sites “enough flexibility to express our terms and conditions on access and use of content,” said Angela Mills Wade, executive director of the European Publishers Council, one of the groups behind the proposal. “That is not surprising. It was invented in the 1990s and things move on.”
Robots.txt was developed in 1994 following concerns that some crawlers were taxing Web sites by visiting them repeatedly or rapidly. Although the system has never been sanctioned by any standards body, major search engines have voluntarily complied.
As search engines expanded to offer services for displaying news and scanning printed books, news organizations and book publishers began to complain.
The proposed extensions, known as Automated Content Access Protocol, partly grew out of those disputes. Leading the ACAP effort were groups representing publishers of newspapers, magazines,…