On Wednesday, Google announced plans to acquire a startup that helps Web sites combat spam and fraud. Google is investing an undisclosed amount to bring reCAPTCHA into its technology fold to address scanning challenges in the Google Books project.
reCAPTCHA is a free anti-bot service that helps digitize books. The company also provides CAPTCHAs to help protect more than 100,000 Web sites. A CAPTCHA is a program that can detect whether its user is a human or a computer.
CAPTCHAs appear as images with distorted text at the bottom of Web registration forms and are used by many Web sites to prevent abuse from automated programs written to generate spam. But Google sees it as a way to teach computers to read.
Teaching Computers to Read
Luis von Ahn, cofounder of reCAPTCHA, and Google product manager Will Cathcart explained the reCAPTCHA twist: The words in many of the CAPTCHAs provided by reCAPTCHA come from scanned archival newspapers and old books.
“Computers find it hard to recognize these words because the ink and paper have degraded over time, but by typing them in as a CAPTCHA, crowds teach computers to read the scanned text,” von Ahn and Cathcart explained. “In this way, reCAPTCHA’s unique technology improves the process that converts scanned images into plain text, known as optical character recognition (OCR).”
Now here’s the Google-reCAPTCHA connection: OCR also powers large-scale text-scanning projects like Google Books and Google News Archive Search. As Google sees it, having the text version of documents is important because plain text can be searched, easily rendered on mobile devices, and displayed to visually impaired users.
Google plans to apply the reCAPTCHA technology not only to increase fraud and spam protection for Google products but also to improve the books and newspaper scanning process. Google will also continue to allow Web-site owners…