Open Access NewsNews from the open access movement Jump to navigation |
|||
Google will share trillion-word dataset
Google has been using its huge index for internal research, but has decided to release a version of it to benefit the research of others. It won't be online for downloading because nobody could download a one trillion word dataset. But it will be available on DVDs, apparently at cost. From Thursday's announcement:
Here at Google Research we have been using word n-gram models for a variety of R&D projects, such as statistical machine translation, speech recognition, spelling correction, entity detection, information extraction, and others....We found that there's no data like more data, and scaled up the size of our data by one order of magnitude, and then another, and then one more - resulting in a training corpus of one trillion words from public Web pages. |
|||