Open Access News

News from the open access movement


Wednesday, June 21, 2006

The data deluge --preserving it, organizing it, accessing it

Scott Carlson, Lost in a Sea of Science Data, Chronicle of Higher Education, June 23, 2006. Excerpt:
Science is experiencing revolutionary changes thanks to digital technology, with computers generating a flood of valuable data for scientists to interpret. But that flood could drown science. Data from experiments conducted as recently as six months ago might be suddenly deemed important, but researchers might never find those numbers — or if they did, might not know what the numbers meant. Lost in some research assistant's computer, the data are often irretrievable or an indecipherable string of digits....To vet experiments, correct errors, or find new breakthroughs, scientists desperately need better ways to store and retrieve research data, says [James M. Caruthers, a professor of chemical engineering at Purdue University], or "we are going to be more and more inefficient in the science that we do in the future."...

The concept of creating shared archives of raw data in those fields is still relatively new, and scientists are still debating who should do the storage — or whether data should be shared at all. But a few colleges and universities have begun experimenting with data libraries. Those projects unite people who don't usually work together, says Clifford A. Lynch, director of the Coalition for Networked Information. "Scientists and scholars on one side and library and IT folks on the other are all feeling their way for the right roles for everybody....The big thing hanging over all of this is funding," he says, adding that agencies like the National Science Foundation are accustomed to supporting science projects and experiments, not infrastructure, like centralized archives....

Purdue librarians were encouraged to tackle the challenge of science data by James L. Mullins, dean of libraries there. Mr. Mullins had worked at the Massachusetts Institute of Technology, where various archiving projects are under way as part of a well-known project called DSpace. Although librarians are working with scientists and technology staff members to apply for grants, the archiving project for now is supported almost entirely by Purdue. Purdue's data-repository model defies some traditional conceptions of an archive. The data will not be stored in a central location on campus, like books stored in the stacks. Instead, the data will reside on the hard drives of faculty members, on departmental servers, or on the TeraGrid, a large-scale computing project run by a handful of institutions, including Purdue D. Scott Brandt, an associate dean at Purdue's library who is directing the project, describes it as a "distributed institutional repository."...

Mr. Caruthers has worked closely with industry and, like many scientists, has worked with data that hold secrets to key discoveries. If data are stored in an archive, researchers should be allowed to keep all or part of it secret, he says. Except for the metadata, Mr. Brandt interjects. "The data that describes what your data is and what it can do would be public — or should be public," he says. Mr. Caruthers shrugs skeptically. Even the simplest metadata describe what a researcher is working on, and can provide advantages to competitors.

Tomorrow (June 22) at 2:00 pm U.S. Eastern time, the Chronicle will host a live online colloquy on these issues with D. Scott Brandt, associate dean for research at Purdue University's libraries. If you can't participate, the Chronicle will publish a transcript later.

Update (6/23/06). The transcript is now online. Excerpt:

Question from Pamela Alexander, University of Pennsylvania:
While an obviously important part of the problem is technical, another aspect is the need for establishing policy for sharing data. As was described in this article, many researchers are reluctant to lose a perceived competitive edge by making their data available to others. However, if this goal is seen as desirable, it will be necessary for federal granting agencies to develop incentives and even requirements for researchers to archive data from funded studies and to provide detailed metadata to make these data accessible to others. As some research agencies (such as the National Institute of Justice) have done, providing funding for secondary data analysis is also another piece of the solution.

D. Scott Brandt:
Good comment --this aspect is important to researchers. One thing is that we think it is critical to have institutional commitment, and we have written this into our strategic plan. Also, at the federal and funding agency level, there is movement to make research results, and in some cases data, available as a stipulation of accepting the grant. For instance, the NIH says it “strongly encourages” pubic accessibility and this is widely interpreted as precursor to further funding. Also, the Cornyn-Lieberman legislation passed recently was intended to ensure research is available to the public (this covers 11 agencies).