Open Access News

News from the open access movement


Wednesday, December 20, 2006

An author addendum for open data

Bill  Hooker, Where are the data? Can I have them? What can I do with them? Open Reading Frame, December 17, 2006.  This is a short excerpt from a long lost.  For critical detail, you should read the whole thing.

...Now, 12 years [after Stevan Harnad's subversive proposal], Open Access is gathering momentum and forward-looking advocates of knowledge as a public good are thinking about Open Data (some extra background here).  Peter Murray-Rust recently stepped up with a subversive proposal of his own:

The simplest thing that researchers can do [to promote Open Data] is to add a Creative Commons license to their data....

I think Peter's proposal is a good one, similar in form and effect to the SPARC author addendum.  Importantly, Science Commons also offers author addenda, and will soon offer them in the machine-, human- and lawyer-readable versions that come with all Creative Commons licenses; as Peter notes, the machine-readable version is crucial to full Open Data utility.  Use of the proposed Open Data addendum (in combination, where necessary, with an Open Access addendum) would clarify the legal status of an author's data, provided we get the wording right.  Herewith some thoughts on how to do that, based on the questions in the title.

First, note that papers do not usually contain raw (useful, useable) data. They contain, say, graphs made from such data, or bitmapped images of it -- as Peter says, the paper offers hamburger when what we want is the original cow....

So if authors want to make their data openly and usefully available, they will need to host it themselves or find someone to host it for them.  Many journals will host supplementary information, and many institutional repositories will take datasets as well as manuscripts.  I have been saying for some time that it should by now be de rigueur to make one's raw data available with each publication. This is very rarely done -- even supplementary information, when I have come across it, tends to be of the hamburger-rather-than-cow variety and so not very useful....

Second, there is the issue of licensing ("Can I have them?  What can I do with them?").... 

It's true but it's simply not enough that, having published in BMC, the authors are probably amenable to giving me the data and allowing me to do with them as I please.  I need unfettered access to the data at the same time as I access the paper.  Even as a human I don't have time to chase down permission for every dataset I want to re-use, and if I'm data-mining by web crawler I need machine-readable licenses that tell my robot what it can have....

So how about Public Library of Science and Hindawi, the other major OA publishers?  Well, Hindawi seems to say nothing about data whatsoever, only that authors retain copyright and articles are published under a CC Attribution license.  PLoS also publishes everything under a CC Attribution license, which says nothing about data, but if you dig a bit you find encouraging things in the editorial/publishing policies....

I suspect another toothless tiger.  It's not that I want the tiger to have teeth, that is, for journals to actively police data availability, but that I wonder why I have to go digging around the website just to find this wishy-washy nod in the general direction of Open Data.  To illustrate my point here, suppose I read a paper in PLoS Biology, and I want to get my hands on some raw data from that paper: where are they?  Can I have them?  What can I do with them?  All of these things are, basically, left up to the authors.

Now remember that these highly unsatisfactory examples are drawn from the most prominent Open Access publishing houses, which might be expected to be much more supportive of Open Data than commercial publishers.  Thus the power of Peter's Open Data addendum becomes apparent: it is attached directly to the paper, so readers do not have to go hunting through journal websites to find out the intellectual property status and location of interesting datasets.  It allows authors to take control....

So, finally, let me take a stab at a draft Open Data addendum....

That's not perfect, not by a long shot -- most especially not for automated data mining, which requires machine-readable metadata and data. It should, however, do what Peter suggests: provide some relief from endless rounds of find-the-permissions, and get a much-needed conversation underway.