Ben Zimmer, executive producer of a Web site and software package called the Visual Thesaurus, was seeking the earliest use of the phrase “you’re not the boss of me.” Using a newspaper database, he had found a reference from 1953.
But while using Google’s book search recently, he found the phrase in a short story in “The Church,” a periodical published in 1883 and scanned from the Bodleian Library, at the University of Oxford.
Ever since Google began scanning printed books four years ago, scholars and others with specialized interests have been able to tap a trove of information that had been locked away in libraries and in antiquarian bookstores.
Dan Clancy, the engineering director for Google’s book search, said that every month users view at least 10 pages of more than half of the one million out-of-copyright books that Google has scanned into its servers.
Zimmer, whose site is visualthesaurus.com, said Google’s book search “allows you to look for things that would be very difficult to search for otherwise.”
A settlement in October with authors and publishers who had brought two copyright lawsuits against Google will make it possible for users to read a far greater collection of books, including many still under copyright protection.
The agreement, pending approval by a judge this year, also paves the way for both sides to make profits from digital versions of books.
Just what kind of commercial opportunity the settlement represents is unknown, but few expect it to generate significant earnings for any individual author. Even Google does not necessarily expect the book program to contribute significantly to its bottom line.
“We did not think necessarily we could make money,” Sergey Brin, co-founder of Google and its president of technology, said during an interview. “We just feel this is part of our core mission. There is fantastic information in books….