This should probably be titled “are research datasets comprised of facts and does it matter?”. It certainly does appear to matter whether datasets are comprised of facts, as in some legal jurisdictions facts are not copyrightable. If this is so, then without other protection such as the EU Database Right, or perhaps contract law, then there is no basis for licences (you can’t control someone else’s use unless you have a right to exercise that control). This is part of the argument that led Science Commons to abandon attempts to find variants of the Creative Commons licences for datasets and databases, in favour of its proposals for putting datasets into the public domain.
This is an appealing solution in some research contexts, but worrying for other kinds of research. This is not for reasons of profit, but of ethics. Many medical, social science, anthropological, financial and other datasets contain data that are private, perhaps personally, culturally or corporately; these data were usually gathered with some kind of informed consent on use, and have to be protected. They cannot be placed into the public domain. If they are to be made available at all for re-use, there must be terms and conditions attached, ie some kind of licence.
I’m not attempting to argue the legal angle here. But I am interested in the “factness” of the data, that might inform the legal angle.
One might assume the height of Mount Everest is a fact. But check out the Wikipedia article on the subject to see a range of results. One might assume that the physical properties of chemical substances are facts, but check out Chemspider’s approach of assembling different measurements with their provenance (see http://www.chemspider.com/blog/there-are-no-facts-in-science-only-measurement-embedded-within-assumptions.html which links to Jen-Claude Bradley’s earlier UsefulChem article ). Or think of a geospatial database, some elements must be pretty much “skill and judgment” rather than facts, such as the point where a river debouches into the sea. Finally, one might assume that the names of the winners of horse races are facts, and so they are, but only after a race committee has adjudicated on the photo-finish, or whether interference took place.
In practice, what goes into datasets is rarely what is directly measured; it is almost always highly derived through various computations, adjustments and combinations. Environmental sciences can be quite explicit on this, see for example the British Atmospheric Data Centre’s description of the UARS (Upper Atmosphere Research Satellite) data levels. Here level 0 is the raw output data streaming from telemetry and instrumentation, effectively at the level of voltage changes; it is devoid of context. Level 1 data has been converted to the physical properties being measured, but will still be in formats tied to the instrument. Level 2 is post-calibration, and would refer to entities such as calculated geophysical profiles. Level 3 would be gridded and interpolated, and at this level there might be no clear correspondence with any observations (but there should be a clear computational lineage or provenance path linking these steps).
So we seem to be in a situation where datasets contain highly derived data, at some creative distance from direct observations, and what we think of as facts are (or ought to be) contestable consensus based on potentially conflicting evidence,
In fact (hah!) after a while it becomes hard to think of any good example of real science/research data that are facts. The question is, does this matter enough to make any difference?
Showing posts with label Databases. Show all posts
Showing posts with label Databases. Show all posts
Thursday, 26 March 2009
Wednesday, 5 November 2008
Some interesting posts elsewhere
I’m sorry for the gap in posting; I’ve been taking a couple of weeks of leave at the end of my trip to Australia. Since return I’ve been catching up on my blog reading, and there are some interesting posts around.
A couple of people (Robin Rice and Jim Downing in particular) have mentioned the post Modelling and storing a phonetics database inside a store, from the Less Talk, More Code blog (Ben O'Steen). This is a practical report on the steps Ben took to put a database into a Fedora Commons-based repository. He details the analysis he went through, the mappings he made, the approaches to capturing representation information, to making the data citable at different levels of granularity, and an interesting approach that he calls “curation by addition”, which appears to be a way of curating the data incrementally, capturing provenance information of all the changes made. It’s a great report, and I look forward to more practical reports of this nature.
Quite a different post on peanubutter (whose author might be Frank Gibson): The Triumvirate of Scientific Data discusses ideas that he suggests relate to the significant properties of science data. His triumvirate comprises
To me, there seemed to be strong resonances between his argument and some of the OAIS concepts, particularly Representation Information. However, context, syntax and semantics might be a more approachable set of labels than RepInfo!
A couple of people (Robin Rice and Jim Downing in particular) have mentioned the post Modelling and storing a phonetics database inside a store, from the Less Talk, More Code blog (Ben O'Steen). This is a practical report on the steps Ben took to put a database into a Fedora Commons-based repository. He details the analysis he went through, the mappings he made, the approaches to capturing representation information, to making the data citable at different levels of granularity, and an interesting approach that he calls “curation by addition”, which appears to be a way of curating the data incrementally, capturing provenance information of all the changes made. It’s a great report, and I look forward to more practical reports of this nature.
Quite a different post on peanubutter (whose author might be Frank Gibson): The Triumvirate of Scientific Data discusses ideas that he suggests relate to the significant properties of science data. His triumvirate comprises
"content, syntax, and semantics, or more simply put -What do we want to say? How do we say it? What does it all mean?"Oddly, the discussion associated with this blog post is on Friendfeed rather than associated with the blog itself. Very interesting to see the discussion recorded like that, and in the process see at least one sceptic become more convinced!
To me, there seemed to be strong resonances between his argument and some of the OAIS concepts, particularly Representation Information. However, context, syntax and semantics might be a more approachable set of labels than RepInfo!
Labels:
Citation,
Curation,
Databases,
OAIS,
Provenance,
Repositories,
Representation Information
Tuesday, 22 July 2008
How open is that data?
Thanks to the Science Commons blog for drawing this article on Nature Precedings to my attention:
DE ROSNAY, M. D. (2008) Check Your Data Freedom: A Taxonomy to Assess Life Science Database Openness. Nature Precedings. doi:10.1038/npre.2008.2083.1
There is an interesting discussion of the issues affecting accessibility (like "no policy visible on the web site"), and the article includes this good set of questions any data curators should ask themselves:
DE ROSNAY, M. D. (2008) Check Your Data Freedom: A Taxonomy to Assess Life Science Database Openness. Nature Precedings. doi:10.1038/npre.2008.2083.1
There is an interesting discussion of the issues affecting accessibility (like "no policy visible on the web site"), and the article includes this good set of questions any data curators should ask themselves:
"A. Check your database technical accessibilityAll good questions to ask, although in some fields the right answer (for some datasets) to questions such as B.4 should be "No" (ethical considerations, for example, might dictate otherwise).
A.1. Do you provide a link to download the whole database?
A.2. Is the dataset available in at least one standard format?
A.3. Do you provide comments and annotations fields allowing users to understand the data?
B. Check your database legal accessibility
B.1. Do you provide a policy expressing terms of use of your database?
B.2. Is the policy clearly indicated on your website?
B.3. Are the terms short and easy to understand by non-lawyers?
B.4. Does the policy authorize redistribution, reuse and modification without restrictions or contractual requirements on the user or the usage?
B.5. Is the attribution requirement at most as strong as the acknowledgment norms of your scientific community?"
Labels:
Data sharing,
Databases,
Digital Curation,
Licensing,
Science Commons
Monday, 7 July 2008
Middle Earth survey preservation
This isn't, as the title of the post might suggest, a posting about preserving geological or geographical data. The CADAIR repository at Aberystwyth University recently ingested a unusually large audience response survey to the Lord of the Rings film, containing just under 25,000 responses from speakers of 14 different languages. One of the repository workers, Stuart Lewis, has posted the details of this on his blog. In short, two versions of the database have been archived: an MS Access version, and an XML representation created by MS Access. These are accompanied by PDF and DOC versions of a user guide to the LotR survey data, and they are all accessible via the local CADAIR DSpace installation.
It's not clear how long the repository will preserve these items for and I couldn't find a preservation policy to provide any insight on this. We had a quick look at both versions and spotted a few issues that affect usability regardless of intended length of retention, such as:
It's not clear how long the repository will preserve these items for and I couldn't find a preservation policy to provide any insight on this. We had a quick look at both versions and spotted a few issues that affect usability regardless of intended length of retention, such as:
- Several questions contain numeric values as answers, but it's not clear what the values are supposed to represent. Eg, on the gender question, does 1 = male or female? You could probably work them out by cross referencing the data against the copy of the questionnaire included in the accompanying user guide, but this isn't a failsafe technique!
- You still need the user guide to understand the content, even for descriptive plain text entries, as column names are only abstracts of the question. For example, it's not clear what the answers in the column 'middle earth' are actually addressing without recourse to the user guide. This is the same in both the XML and MS Access version and highlights a potential issue in relying on the built-in converter tool alone because it doesn't allow you to manipulate the XML and add extra value. For instance, why not include more detailed information about the questions in the XML file?
- It would be useful to know from the dataset record entry in CADAIR that it contains text in multiple languages - the dc.language value is english and this is correct insofar as field names go, but not for all of the content.
Wednesday, 25 June 2008
4th International Digital Curation Conference - Radical Sharing: Transforming Science?
This is a reminder that closing date for submission of full papers, posters and demos for the conference (details below) is 25 July 2008 . We invite submissons from individuals, organisations and institutions across all disciplines and domains engaged in the creation, use and management of digital data, especially those involved in the challenge of curating data in e-Science and e-Research. Templates for submissions and other details are available at http://www.dcc.ac.uk/events/dcc-2008/ .
The 4th International Digital Curation Conference will be held on 1-3 December 2008 at the Hilton Grosvenor Hotel, Edinburgh, Scotland, UK. The conference will be held in partnership with the National e-Science Centre and supported by the Coalition for Networked Information (CNI).
We are sure there are plenty of you out there with great ideas to share with colleagues; still time to get writing!
The 4th International Digital Curation Conference will be held on 1-3 December 2008 at the Hilton Grosvenor Hotel, Edinburgh, Scotland, UK. The conference will be held in partnership with the National e-Science Centre and supported by the Coalition for Networked Information (CNI).
We are sure there are plenty of you out there with great ideas to share with colleagues; still time to get writing!
Labels:
Conference,
Data,
Data Re-use,
Data Services,
Data sharing,
Databases,
Digital Curation,
IDCC4
Friday, 25 April 2008
A Thousand Open Molecular Biology Databases
In January of each year, Nucleic Acids Research (NAR) publishes a special issue on databases for molecular biology research. To be considered, databases have to be open access (they specifically mean browsable without a username, password or payment, although it is possible there are conditions). The staggering thing is, in the past year the number of such databases passed 1,000!
My institution has a subscription to NAR, but the nice thing about the Database Issue is that it has itself been Open Access for the past 4 years or so. So you can check it out for yourself if you are interested.
I thought it might be interesting to go back and trace how they managed to get to 1,000+ databases in 15 years or so. It turns out to be relatively easy to check, back to about the year 2001; prior to that, as far as I can tell, you have to count them yourself. I did the count for 1999, so here’s a little picture of the growth since then (see Burks, 1999):

It is quite a staggering growth.
In recent years, the compilation article has been written by Michael Y. Galperin of the National Center for Biotechnology Information, US National Library of Medicine. He reports in the 2008 article that the complete list and summaries of 1078 databases are available online at the Nucleic Acids Research web site, http://www.oxfordjournals.org/nar/database/a.
This year’s article has some interesting comments on databases; I particularly liked this one on Deja Vu, which uses a tool called eTBLAST “to find highly similar abstracts in bibliographic databases, including MEDLINE…. Some highly similar publications, however, come from different authors and look extremely suspicious” (Galperin, 2008). Curious! As usual he also reports on databases that appear to be no longer maintained and have been dropped from the list (around two dozen this year). Sometimes this is related to (perhaps because of, or maybe the cause of) the content being available in other databases.
There has in fact been relatively little attrition; they claim not to re-use accession numbers, and the highest accession number so far is 1176, implying that just 98 databases have been dropped from the list! Galperin suggests “that the databases that offer useful content usually manage to survive, even if they have to change their funding scheme or migrate from one host institution to another. This means that the open database movement is here to stay, and more and more people in the community (as well as in the financing bodies) now appreciate the importance of open databases in spreading knowledge. It is worth noting that the majority of database authors and curators receive little or no remuneration for their efforts and that it is still difficult to obtain money for creating and maintaining a biological database. However, disk space is relatively cheap these days and database maintenance tools are fairly straightforward, so that a decent database can be created on a shoestring budget, often by a graduate student or as a result of a postdoctoral project. […] Subsequent maintenance and further development of these databases, however, require a commitment that can only be applauded.” (Galperin, 2005)
So is this vanity databasing (this from a blog author, mind)? “In the very beginning of the genome sequencing era, Walter Gilbert and colleagues warned of 'database explosion', stemming from the exponentially increasing amount of incoming DNA sequence and the unavoidable errors it contains. Luckily, this threat has not materialized so far, due to the corresponding growth in computational power and storage capacity and the strict requirements for sequence accuracy.” (Galperin, 2004)
It’s not clear from the quote above how worth while or well-used the databases are. In the 2006 article, Galperin began looking at measures of impact, using the Science Citation Index. We can see from his reference list that he expects the NAR paper to stand proxy for the database. The highest cited were Pfam, GO, UniProt, SMART and KEGG, all highly used “instant classics” (with >100 citations each in 2 years!). However, he writes: “On the other side of the spectrum are the databases that have never been cited in these 2 years, even by their own authors. This does not mean, of course, that these databases do not offer a useful content but one could always suggest a reason why nobody has used this or that database. Usually these databases were too specific in scope and offered content that could be easily found elsewhere.” (Galperin, 2006)
In the 2007 article, Galperin returned to this issue of how well databases are used. “However, citation data can be biased; e.g. in many articles use of information from publicly available databases is acknowledged by providing their URLs, or not acknowledged at all. Besides, some databases could be cited on the web sites and in new or obscure journals, not covered by the ISI Citation Index.” (Galperin, 2007) He then goes on to describe some alternative measures he has investigated to proxy for this citation problem. This is a real issue, I think; data and dataset citations are not made as often or as consistently as they should be, and advice is often conflicting and itself conflicts with the conflicting standards (of which perhaps the best is the NLM standard). Indeed, the NAR articles describing databases seem to stand proxy for the databases: “the user typically starts by finding a database of interest in PubMed or some other bibliographic database, then proceeds to browse the full text in the HTML format. If the paper is interesting enough, s/he would download its text in the PDF format. Finally, if the database turns to be useful, it might be acknowledged with a formal citation.”
This is probably enough for one blog post, but I’ll return, I think, to have a look at some of these databases in a bit more detail.
BURKS, C. (1999) Molecular Biology Database List. Nucl. Acids Res., 27, 1-9. http://nar.oxfordjournals.org/cgi/content/abstract/27/1/1
GALPERIN, M. Y. (2004) The Molecular Biology Database Collection: 2004 update. Nucl. Acids Res., 32, D3-22. http://nar.oxfordjournals.org/cgi/content/abstract/32/suppl_1/D3
GALPERIN, M. Y. (2005) The Molecular Biology Database Collection: 2005 update. Nucleic Acids Research, 33.
GALPERIN, M. Y. (2006) The Molecular Biology Database Collection: 2006 update
. Nucleic Acids Research, 34. http://nar.oxfordjournals.org/cgi/content/full/34/suppl_1/D3
GALPERIN, M. Y. (2007) The Molecular Biology Database Collection: 2007 update. Nucleic Acids Research, 35, D3-D4. http://nar.oxfordjournals.org/cgi/content/abstract/35/suppl_1/D3
GALPERIN, M. Y. (2008) The Molecular Biology Database Collection: 2008 update. Nucleic Acids Research, 36, D2-D4. http://dx.doi.org/10.1093/nar/gkm1037
My institution has a subscription to NAR, but the nice thing about the Database Issue is that it has itself been Open Access for the past 4 years or so. So you can check it out for yourself if you are interested.
I thought it might be interesting to go back and trace how they managed to get to 1,000+ databases in 15 years or so. It turns out to be relatively easy to check, back to about the year 2001; prior to that, as far as I can tell, you have to count them yourself. I did the count for 1999, so here’s a little picture of the growth since then (see Burks, 1999):

It is quite a staggering growth.
In recent years, the compilation article has been written by Michael Y. Galperin of the National Center for Biotechnology Information, US National Library of Medicine. He reports in the 2008 article that the complete list and summaries of 1078 databases are available online at the Nucleic Acids Research web site, http://www.oxfordjournals.org/nar/database/a.
This year’s article has some interesting comments on databases; I particularly liked this one on Deja Vu, which uses a tool called eTBLAST “to find highly similar abstracts in bibliographic databases, including MEDLINE…. Some highly similar publications, however, come from different authors and look extremely suspicious” (Galperin, 2008). Curious! As usual he also reports on databases that appear to be no longer maintained and have been dropped from the list (around two dozen this year). Sometimes this is related to (perhaps because of, or maybe the cause of) the content being available in other databases.
There has in fact been relatively little attrition; they claim not to re-use accession numbers, and the highest accession number so far is 1176, implying that just 98 databases have been dropped from the list! Galperin suggests “that the databases that offer useful content usually manage to survive, even if they have to change their funding scheme or migrate from one host institution to another. This means that the open database movement is here to stay, and more and more people in the community (as well as in the financing bodies) now appreciate the importance of open databases in spreading knowledge. It is worth noting that the majority of database authors and curators receive little or no remuneration for their efforts and that it is still difficult to obtain money for creating and maintaining a biological database. However, disk space is relatively cheap these days and database maintenance tools are fairly straightforward, so that a decent database can be created on a shoestring budget, often by a graduate student or as a result of a postdoctoral project. […] Subsequent maintenance and further development of these databases, however, require a commitment that can only be applauded.” (Galperin, 2005)
So is this vanity databasing (this from a blog author, mind)? “In the very beginning of the genome sequencing era, Walter Gilbert and colleagues warned of 'database explosion', stemming from the exponentially increasing amount of incoming DNA sequence and the unavoidable errors it contains. Luckily, this threat has not materialized so far, due to the corresponding growth in computational power and storage capacity and the strict requirements for sequence accuracy.” (Galperin, 2004)
It’s not clear from the quote above how worth while or well-used the databases are. In the 2006 article, Galperin began looking at measures of impact, using the Science Citation Index. We can see from his reference list that he expects the NAR paper to stand proxy for the database. The highest cited were Pfam, GO, UniProt, SMART and KEGG, all highly used “instant classics” (with >100 citations each in 2 years!). However, he writes: “On the other side of the spectrum are the databases that have never been cited in these 2 years, even by their own authors. This does not mean, of course, that these databases do not offer a useful content but one could always suggest a reason why nobody has used this or that database. Usually these databases were too specific in scope and offered content that could be easily found elsewhere.” (Galperin, 2006)
In the 2007 article, Galperin returned to this issue of how well databases are used. “However, citation data can be biased; e.g. in many articles use of information from publicly available databases is acknowledged by providing their URLs, or not acknowledged at all. Besides, some databases could be cited on the web sites and in new or obscure journals, not covered by the ISI Citation Index.” (Galperin, 2007) He then goes on to describe some alternative measures he has investigated to proxy for this citation problem. This is a real issue, I think; data and dataset citations are not made as often or as consistently as they should be, and advice is often conflicting and itself conflicts with the conflicting standards (of which perhaps the best is the NLM standard). Indeed, the NAR articles describing databases seem to stand proxy for the databases: “the user typically starts by finding a database of interest in PubMed or some other bibliographic database, then proceeds to browse the full text in the HTML format. If the paper is interesting enough, s/he would download its text in the PDF format. Finally, if the database turns to be useful, it might be acknowledged with a formal citation.”
This is probably enough for one blog post, but I’ll return, I think, to have a look at some of these databases in a bit more detail.
BURKS, C. (1999) Molecular Biology Database List. Nucl. Acids Res., 27, 1-9. http://nar.oxfordjournals.org/cgi/content/abstract/27/1/1
GALPERIN, M. Y. (2004) The Molecular Biology Database Collection: 2004 update. Nucl. Acids Res., 32, D3-22. http://nar.oxfordjournals.org/cgi/content/abstract/32/suppl_1/D3
GALPERIN, M. Y. (2005) The Molecular Biology Database Collection: 2005 update. Nucleic Acids Research, 33.
GALPERIN, M. Y. (2006) The Molecular Biology Database Collection: 2006 update
. Nucleic Acids Research, 34. http://nar.oxfordjournals.org/cgi/content/full/34/suppl_1/D3
GALPERIN, M. Y. (2007) The Molecular Biology Database Collection: 2007 update. Nucleic Acids Research, 35, D3-D4. http://nar.oxfordjournals.org/cgi/content/abstract/35/suppl_1/D3
GALPERIN, M. Y. (2008) The Molecular Biology Database Collection: 2008 update. Nucleic Acids Research, 36, D2-D4. http://dx.doi.org/10.1093/nar/gkm1037
Labels:
Biocuration,
Citation,
Data,
Databases,
Digital Curation
Monday, 31 March 2008
UK Repositories claiming to hold data
The OpenDOAR and ROAR services both present self-reported claims by repositories across the world about their contents, backed up by some harvested facts. I’m interested in those UK repositories that claim to hold data.
My first problem is that neither repository allows me simply to choose data. OpenDOAR allows me to search on “Datasets” (63 world-wide, 8 in the UK), while ROAR allows me to search for “Database/A&I Index” (24 world-wide, 6 in the UK). I thought the latter was a surprisingly “library science” classification, given the origins of ROAR. Not surprisingly, most repositories are in only one of the lists. Also not surprisingly given the origins of these services in the Open Access and OAI-PMH movements, there are many first class data repositories NOT listed here (UKDA and BADC, for example).
The UK repositories listed are:
OpenDOAR “Datasets”
The 3 that do have serious amounts of data are DSpace @ Cambridge, eCrystals and NDAD. DSpace @ Cambridge is dominated by the 100,000 ++ collection of chemical structures encoded in CML, but there are plenty of other datasets there, including some from Archaeology. Sadly, there are plenty of empty collections, and many collections where the last deposit was 2006 (I guess around when the funded project died). eCrystals is completely crystal structures, and has some very nice features; find a compound, and as you look perhaps rather bemused at the page, a Java object loads and there you have a rotatable image of the molecular structure before your eyes on the data page! NDAD also has many ex-Government datasets, some of them very large.
ROAR “Database/A&I Index”
It’s a rather sad study! I do hope that the Open Repositories 2008 conference in Southampton over the next couple of days leads to an improvement. I can't get there, unfortunately, but I hope someone will report from it here. I particularly liked the idea of the developers challenges. Can we have some oriented to data, please?
My first problem is that neither repository allows me simply to choose data. OpenDOAR allows me to search on “Datasets” (63 world-wide, 8 in the UK), while ROAR allows me to search for “Database/A&I Index” (24 world-wide, 6 in the UK). I thought the latter was a surprisingly “library science” classification, given the origins of ROAR. Not surprisingly, most repositories are in only one of the lists. Also not surprisingly given the origins of these services in the Open Access and OAI-PMH movements, there are many first class data repositories NOT listed here (UKDA and BADC, for example).
The UK repositories listed are:
OpenDOAR “Datasets”
- Bristol Repository of Scholarly Eprints (ROSE)
- DSpace @ Cambridge
- eCrystals - Southampton
- Edinburgh DataShare
- Edinburgh Research Archive (ERA)
- Leicester Research Archive (LRA)
- NDAD (National Digital Archive of Datasets)
- Nature Precedings
The 3 that do have serious amounts of data are DSpace @ Cambridge, eCrystals and NDAD. DSpace @ Cambridge is dominated by the 100,000 ++ collection of chemical structures encoded in CML, but there are plenty of other datasets there, including some from Archaeology. Sadly, there are plenty of empty collections, and many collections where the last deposit was 2006 (I guess around when the funded project died). eCrystals is completely crystal structures, and has some very nice features; find a compound, and as you look perhaps rather bemused at the page, a Java object loads and there you have a rotatable image of the molecular structure before your eyes on the data page! NDAD also has many ex-Government datasets, some of them very large.
ROAR “Database/A&I Index”
- Higher Education Empirical Research Database (1642 records)
- NDAD - UK National Digital Archive of Datasets (66 records)
- ReOrient Knowledge Base
- Research Findings Register (1496 records)
- Southampton Crystal Structure Report Archive (165 records)
- The Linnean Collections (14275 records)
It’s a rather sad study! I do hope that the Open Repositories 2008 conference in Southampton over the next couple of days leads to an improvement. I can't get there, unfortunately, but I hope someone will report from it here. I particularly liked the idea of the developers challenges. Can we have some oriented to data, please?
Labels:
Conference,
Data,
Data Services,
Databases,
Open Access,
Repositories,
Research data
Friday, 20 July 2007
Open Data Licensing: is your data safe?
Over on the Nodalities blog, Rob Styles wrote about some of the aspects of open data licensing, and the tricky questions of copyright versus database right. OK, yawn. Let me put that another way… over on the Nodalities blog, Rob Styles writes about whether you can make your data openly accessible on the web without getting totally ripped off in the process. A bit less of a yawn?
One key quote:
The problem is that there is doubt… OK, more than doubt… whether and/or how Copyright applies to databases. And if Copyright does not apply, you don’t get the exclusive control which allows you to apply a conditional licence like Creative Commons. Just to explore a bit further...
Science Commons was set up to look at helping make science data more openly available. But if you look at their FAQ, you can see some real concerns. They pick out several aspects of a database that might be subject to Copyright, including the structure, but also say:
As I’ve said before, I’m not a lawyer. Can a data-oriented lawyer comment?
One key quote:
“Without appropriate protection of intellectual property we have only two extreme positions available: locked down with passwords and other technical means; or wide open and in the public-domain. Polarising the possibilities for data into these two extremes makes opening up an all or nothing decision for the creator of a database.It’s true: to put any conditions over the use of our data, we have to have an exclusive right to control it. Copyright gives its owner that right for a text. If I own the Copyright for my works, I can (and try to) put a Creative Commons licence on it, to allow others to use it but to ask them to give me attribution if they do so.
With only technical and contractual mechanisms for protecting data, creators of databases can only publish them in situations where the technical barriers can be maintained and contractual obligations can be enforced.”
The problem is that there is doubt… OK, more than doubt… whether and/or how Copyright applies to databases. And if Copyright does not apply, you don’t get the exclusive control which allows you to apply a conditional licence like Creative Commons. Just to explore a bit further...
Science Commons was set up to look at helping make science data more openly available. But if you look at their FAQ, you can see some real concerns. They pick out several aspects of a database that might be subject to Copyright, including the structure, but also say:
"In the United States, data will be protected by copyright only if they express creativity. Some databases will satisfy this condition, such as a database containing poetry or a wiki containing prose. Many databases, however, contain factual information that may have taken a great deal of effort to gather, such as the results of a series of complicated and creative experiments. Nonetheless, that information is not protected by copyright and cannot be licensed under the terms of a Creative Commons license."In a note to me Mags McGinley, our legal officer, re-inforces this, and adds:
"Copyright definitely applies to certain elements of a database. Copyright exists in the structure of a database if, by reason of the selection and arrangement, it constitutes the authors own intellectual creation. In addition the contents of database, depending on what they are, may attract their own copyright protection (a simple example might be a database of poems)."But is there a glimmer of hope? The Science Commons FAQ goes on to say:
"Note - for databases subject to the laws of members of the European Union and certain other countries, the law supplies a special right for databases. Except in the Netherlands and Belgium Creative Commons Licenses, Creative Commons licenses do not apply to this right..."Rob Styles also reminds us that in Europe we have this other right: “the EU adopted a robust database right in 1996 while the US ruled against such protection in 1991”.
“Database right in the EU is like Copyright. It is a monopoly, but only on that particular aggregation of the data. The underlying facts are still not protected and there is nothing to stop a second entrant from collecting them independently.”Charlotte Waelde has written a report for the JISC-funded GRADE project on rights that apply to data in geospatial databases. She concluded that Database Copyright does not apply, but the Database Right does apply. She also concluded (my emphasis):
"• Unauthorised taking and making available of substantial parts of the contents of the database will infringe the right of extraction and re-utilisation"and...
"• A lawful user of the database (e.g. the researcher or teacher in an educational institution) may not be prevented from extracting and re-utilising an insubstantial part of the contents of a database for any purposes whatsoever.I am not a lawyer and (try as I might) I couldn't get all the nuances of what she is trying to say, particularly in the last sentence above; however Mags tells me
• A researcher or teacher may not be prevented from extracting a substantial part of the contents of the database for the purposes of non-commercial research or illustration for teaching so long as the source is indicated. Re-utilisation may only be enjoined if the output contains a substantial part of the contents of the protected database"
"The thing there is that there is a difference between extraction and reutilisation which are the two activities that can be prevented by the database right. The fair dealing exceptions for the database right are not as wide as those of copyright and are for some reason limited to the act of extraction."
"So Charlotte is highlighting the maximum you could do in such case where your activities fall within the research/teaching area. This is: extract a substantial part. And then reutilise an insubstantial part (because the database right only limits what you do with substantial parts of the database)."Rob goes on to end his blog entry, saying of rights:
“They allow inventors to disclose their inventions when they might otherwise have had to keep them secret... That's why we've invested in a license to do this, properly, clearly and in a way that stays Open.”He is referring to the Talis Community Licence, which attempts to base a conditional open licence on the Database Right. Trust me, I REALLY want this sort of thing to work, but I worry that the Database Right may not be sufficient as underlying protection to make this licence firm. And what would be the law applying to access FROM a jurisdiction like the US that did not have a Database Right?
As I’ve said before, I’m not a lawyer. Can a data-oriented lawyer comment?
Subscribe to:
Posts (Atom)
