Open Access News

News from the open access movement


Tuesday, December 14, 2004

Google scanning, indexing library books

John Markoff and Edward Wyatt, Google Is Adding Major Libraries to Its Database, New York Times, December 14, 2004. Excerpt: 'Google, the operator of the world's most popular Internet search service, plans to announce an agreement today with some of the nation's leading research libraries and Oxford University to begin converting their holdings into digital files that would be freely searchable over the Web....Google - newly wealthy from its stock offering last summer - has agreed to underwrite the projects being announced today while also adding its own technical abilities to the task of scanning and digitizing tens of thousands of pages a day at each library....Within two decades, most of the world's knowledge will be digitized and available, one hopes for free reading on the Internet, just as there is free reading in libraries today," said Michael A. Keller, Stanford University's head librarian....Last night the Library of Congress and a group of international libraries from the United States, Canada, Egypt, China and the Netherlands announced a plan to create a publicly available digital archive of one million books on the Internet. The group said it planned to have 70,000 volumes online by next April. "Having the great libraries at your fingertips allows us to build on and create great works based on the work of others," said Brewster Kahle, founder and president of the Internet Archive, a San Francisco-based digital library that is also trying to digitize existing print information. The agreements to be announced today will allow Google to publish the full text of only those library books old enough to no longer be under copyright. For copyrighted works, Google would scan in the entire text, but make only short excerpts available online. Each agreement with a library is slightly different. Google plans to digitize nearly all the eight million books in Stanford's collection and the seven million at Michigan. The Harvard project will initially be limited to only about 40,000 volumes. The scanning at Bodleian Library at Oxford will be limited to an unspecified number of books published before 1900, while the New York Public Library project will involve fragile material not under copyright that library officials said would be of interest primarily to scholars.'