There's a concept in maths called "closed but unbounded". I'm not sure it's exactly to the point (I hope that's a pun), but "subjects" seem a bit like that. You can be pretty sure about most of the stuff that's not in a subject (or "domain"), and most of the stuff that is in it, but you can be very puzzled about some of the edges, and can find yourself in some extremely surprising discussions at times about parts of subjects that challenge most of the ideas you had. So subjects turn out to be very un-bounded. (They also tend to fracture, productively.) Perhaps not surprisingly, subjects don't tend to have assets, bank balances, etc. You might say, in those senses, subjects don't exist! They do nevertheless have very real approaches, common standards, ontologies, methods, vocabularies, literatures... and passionate adherents spread across institutions.
Institutions on the other hand, or at least universities, tend to be very material. They do have assets, bank balances, policies, libraries, employees, continuity on a significant scale, even (in the US at least) endowments. They have temporal stability and mass. They collect scholars and scientists in various domains... even if the scientists give their loyalty to their subjects, and are held together only by salaries and a common loathing of the university car parking policy!
Institutions have continuity, and they have libraries, and archives, which in serious ways express that continuity. Libraries are not about print. Libraries are now squarely about knowledge and information expressed in data, whether they know it or not. And the continuity of valuable data is an important reason for libraries to be involved.
But institutions are generic, and libraries are generic, even in more focused institutions like MIT. The library, the archive, the IR, in different ways, are about collecting elements of the scholarly discourse that contribute both globally and locally. So institutional repositories are about generic continuity of data, as libraries are about continuity of collections. IRs create value for the institution, even if it is only a small piece of value (like most other individual "collections" in an institution). If you don't play, you aren't in the game. You know data has value, just not which bits. You need to disclose your scholarly assets, across the spectrum; you can feel proud of doing so, and make a case for local benefit at the same time. You are an institution taking part in a global system; the value may be in the network, but you are part of that network.
But the way an IR treats data is necessarily generic; if you get data from chemistry, engineering, social sciences and performing arts into this under-funded but potentially valuable repository, you will do your best but it will necessarily be variants of generic practice, at best.
So back to the "subject"; if there is a data repository here, it is likely staffed by "domain experts", capable of taking on a "community proxy" role. They know their stuff. They will treat their data in domain-specific ways; they will know where to seek out data to complement their collection, they will know how to make connections between different parts. They can describe it appropriately, they can develop standards with their colleagues. They will know how to help their colleague scientists extract maximum value. Some subject repository managers are seriously concerned about the problems for disciplines if institutional repositories expand into the data "space".
What subject repositories don't usually have, is what institutions have: substantial assets, endowments, bank balances, tenured staff. Usually based around multiple project grants, 5-year core funding is a prized goal at a price of cheese-paring funding, and mid-term reviews every second year. Subject repositories don't have assured continuity, temporal mass.
The NSB LLDDC (Long-lived digital data collections) report, and now the NSF CyberInfrastructure strategy, are aimed at this area; they have spotted the fragility of these subject data collections. In the UK we have possibly even more of a patchwork of funding mechanisms than was observed in the LLDDC report. JISC used to be a significant funder of subject repositories, but in recent years has been retrenching from them, while building up massive funding in IRs. AHRC, as we have seen, is pulling back from funding the AHDS.
So what would make this better? I'd like to see a substantive discussion about the roles and funding mechanisms of subject and institutional repositories. In the UK, this would have to involve at least the Research Councils, Wellcome Trust and JISC. (Perhaps looks less likely than it did when I first wrote this.)
Secondly, I'd like to see JISC in the final tranche of its capital funding (here's the circular recently closed) explore the bounds of what's possible with the data provider/ service provider combination (maybe OAI/ORE will address this a little? Maybe not!). And what if curation is detached from the repository? What if data continuity/preservation is separated from the curation service? Do these questions even make sense?
Maybe a system or federation of sustainable IRs internally divided into sets on subject lines (and hence externally aggregatable along those lines), with subject-oriented curation activities picking up on "invisible college" volunteerism might work? Splitting curation into generic and domain elements... Or other notions, pushing the skill out into the network, the federation, but retaining the data where the assets and continuity lie?
[This posting is based on an email I sent to a closed JISC Repositories advisory group some time ago; it seems even more relevant today...]
Friday, 20 July 2007
Open Data Licensing: is your data safe?
Over on the Nodalities blog, Rob Styles wrote about some of the aspects of open data licensing, and the tricky questions of copyright versus database right. OK, yawn. Let me put that another way… over on the Nodalities blog, Rob Styles writes about whether you can make your data openly accessible on the web without getting totally ripped off in the process. A bit less of a yawn?
One key quote:
The problem is that there is doubt… OK, more than doubt… whether and/or how Copyright applies to databases. And if Copyright does not apply, you don’t get the exclusive control which allows you to apply a conditional licence like Creative Commons. Just to explore a bit further...
Science Commons was set up to look at helping make science data more openly available. But if you look at their FAQ, you can see some real concerns. They pick out several aspects of a database that might be subject to Copyright, including the structure, but also say:
As I’ve said before, I’m not a lawyer. Can a data-oriented lawyer comment?
One key quote:
“Without appropriate protection of intellectual property we have only two extreme positions available: locked down with passwords and other technical means; or wide open and in the public-domain. Polarising the possibilities for data into these two extremes makes opening up an all or nothing decision for the creator of a database.It’s true: to put any conditions over the use of our data, we have to have an exclusive right to control it. Copyright gives its owner that right for a text. If I own the Copyright for my works, I can (and try to) put a Creative Commons licence on it, to allow others to use it but to ask them to give me attribution if they do so.
With only technical and contractual mechanisms for protecting data, creators of databases can only publish them in situations where the technical barriers can be maintained and contractual obligations can be enforced.”
The problem is that there is doubt… OK, more than doubt… whether and/or how Copyright applies to databases. And if Copyright does not apply, you don’t get the exclusive control which allows you to apply a conditional licence like Creative Commons. Just to explore a bit further...
Science Commons was set up to look at helping make science data more openly available. But if you look at their FAQ, you can see some real concerns. They pick out several aspects of a database that might be subject to Copyright, including the structure, but also say:
"In the United States, data will be protected by copyright only if they express creativity. Some databases will satisfy this condition, such as a database containing poetry or a wiki containing prose. Many databases, however, contain factual information that may have taken a great deal of effort to gather, such as the results of a series of complicated and creative experiments. Nonetheless, that information is not protected by copyright and cannot be licensed under the terms of a Creative Commons license."In a note to me Mags McGinley, our legal officer, re-inforces this, and adds:
"Copyright definitely applies to certain elements of a database. Copyright exists in the structure of a database if, by reason of the selection and arrangement, it constitutes the authors own intellectual creation. In addition the contents of database, depending on what they are, may attract their own copyright protection (a simple example might be a database of poems)."But is there a glimmer of hope? The Science Commons FAQ goes on to say:
"Note - for databases subject to the laws of members of the European Union and certain other countries, the law supplies a special right for databases. Except in the Netherlands and Belgium Creative Commons Licenses, Creative Commons licenses do not apply to this right..."Rob Styles also reminds us that in Europe we have this other right: “the EU adopted a robust database right in 1996 while the US ruled against such protection in 1991”.
“Database right in the EU is like Copyright. It is a monopoly, but only on that particular aggregation of the data. The underlying facts are still not protected and there is nothing to stop a second entrant from collecting them independently.”Charlotte Waelde has written a report for the JISC-funded GRADE project on rights that apply to data in geospatial databases. She concluded that Database Copyright does not apply, but the Database Right does apply. She also concluded (my emphasis):
"• Unauthorised taking and making available of substantial parts of the contents of the database will infringe the right of extraction and re-utilisation"and...
"• A lawful user of the database (e.g. the researcher or teacher in an educational institution) may not be prevented from extracting and re-utilising an insubstantial part of the contents of a database for any purposes whatsoever.I am not a lawyer and (try as I might) I couldn't get all the nuances of what she is trying to say, particularly in the last sentence above; however Mags tells me
• A researcher or teacher may not be prevented from extracting a substantial part of the contents of the database for the purposes of non-commercial research or illustration for teaching so long as the source is indicated. Re-utilisation may only be enjoined if the output contains a substantial part of the contents of the protected database"
"The thing there is that there is a difference between extraction and reutilisation which are the two activities that can be prevented by the database right. The fair dealing exceptions for the database right are not as wide as those of copyright and are for some reason limited to the act of extraction."
"So Charlotte is highlighting the maximum you could do in such case where your activities fall within the research/teaching area. This is: extract a substantial part. And then reutilise an insubstantial part (because the database right only limits what you do with substantial parts of the database)."Rob goes on to end his blog entry, saying of rights:
“They allow inventors to disclose their inventions when they might otherwise have had to keep them secret... That's why we've invested in a license to do this, properly, clearly and in a way that stays Open.”He is referring to the Talis Community Licence, which attempts to base a conditional open licence on the Database Right. Trust me, I REALLY want this sort of thing to work, but I worry that the Database Right may not be sufficient as underlying protection to make this licence firm. And what would be the law applying to access FROM a jurisdiction like the US that did not have a Database Right?
As I’ve said before, I’m not a lawyer. Can a data-oriented lawyer comment?
Wednesday, 18 July 2007
Arts and Humanities Data Service... next steps?
In an earlier post, I mentioned the decision by the AHRC (and later JISC) to cease funding the
AHDS from March 2008. Since then the AHRC have re-affirmed their decision. On 28 June, the Future Histories of the Moving Image Research Network made public an open letter to the AHRC, to no avail it would appear.
In her response to the announcements, Sheila Anderson (Director of AHDS) wrote:
Explicit in the AHRC's decision was the view that the community is mature enough to manage its own resources. There is doubt in many people's minds about this, but we are effectively stuck with it. So what are the implications? There are implications both for existing collections and for future arts and humanities resources. I would like to spend a few paragraphs thinking about the existing collections.
AHDS is not monolithic; it is comprised of several separate services (I suspect in what follows I may be using historical rather than current names). We already know that AHRC is privileging the Archaeology Data Service (ADS), which will continue to receive some funding (and which has a diverse funding base), so their resources are presumably safe. The History Data Service (HDS) resources are embedded within the UK Data Archive (which has recently received an additional 5 years funding from ESRC); it would presumably cost more to de-accession those resources than to keep preserving them and making them available, so even if HDS can take no more resources, the existing ones should be safe. Literature, Language and Linguistics is closely related to the Oxford Text Archive; I imagine the same kinds of arguments would apply there.
I have heard suggestions that Kings College London might continue to support the AHDS Executive for a period, and it appears there are some discussions with JISC about some kind of support "to ensure the expertise and achievements of the AHDS are not lost to the community".
That leaves Performing Arts and Visual Arts; I can't even surmise what their future might be, since I don't know enough about their funding and local environment.
I appreciate that it's still early days, and no doubt crucial discussions are going on behind the scenes. But if any part of AHDS resources are in danger of loss, the resource owners need to consider plans to deal with those resources in the future. This will take some time, particularly since for more complex resources, it is clear that existing repositories are generally NOT yet adequate for purpose. I guess the picture itself will be complex; I can think of at least these categories:
Are we OK? Is there more? Who knows! I think we need much better tools to tell what is "at risk", so that plans can start being made. Of course, this could be happening, maybe I'm just not "in the loop".
Will AHRC consider bids for funding transitional work? I certainly hope so, although I don't know how this might be done. JISC is (I believe) planning one last round of its Capital Programme. Will they include provisions to enhance repositories so as to take these more complex resources? I certainly hope so!
AHDS from March 2008. Since then the AHRC have re-affirmed their decision. On 28 June, the Future Histories of the Moving Image Research Network made public an open letter to the AHRC, to no avail it would appear.
In her response to the announcements, Sheila Anderson (Director of AHDS) wrote:
"In the meantime, and at least until 31st March 2008, the AHDS will continue to give advice and guidance on all matters relating to the creation of digital content arising from or supporting research, teaching and learning across the arts and humanities, including technical and metadata standards and project management. If you have a data creation project, please do not hesitate to contact us for advice.This is a great expression of commitment, and deserves our support. However, the lack of long-term funding must raise questions of sustainability.
The AHDS will continue to work with those creating important digital resources to advise on the best methods for keeping these valuable resources available and accessible for the long-term in a form that encourages their further use for answering new research questions, and their use in teaching and learning. This advice will include exploring with content creators and owners suitable repositories in which they might deposit their materials for long term curation and preservation, and how to ensure that their materials can continue to be discovered and used by the wider community. If you are currently in negotiation with the AHDS to deposit your digital collection, please continue to work with us to ensure the future sustainability and accessibility of your resource.
The AHDS will continue to make available its rich collection of digital content for use in research, teaching and learning, and to preserve those collections in its care. The AHDS intends to discuss with the JISC and the AHRC the long term future of these collections beyond April 2008 with the intention of securing their continued preservation and availability."
Explicit in the AHRC's decision was the view that the community is mature enough to manage its own resources. There is doubt in many people's minds about this, but we are effectively stuck with it. So what are the implications? There are implications both for existing collections and for future arts and humanities resources. I would like to spend a few paragraphs thinking about the existing collections.
AHDS is not monolithic; it is comprised of several separate services (I suspect in what follows I may be using historical rather than current names). We already know that AHRC is privileging the Archaeology Data Service (ADS), which will continue to receive some funding (and which has a diverse funding base), so their resources are presumably safe. The History Data Service (HDS) resources are embedded within the UK Data Archive (which has recently received an additional 5 years funding from ESRC); it would presumably cost more to de-accession those resources than to keep preserving them and making them available, so even if HDS can take no more resources, the existing ones should be safe. Literature, Language and Linguistics is closely related to the Oxford Text Archive; I imagine the same kinds of arguments would apply there.
I have heard suggestions that Kings College London might continue to support the AHDS Executive for a period, and it appears there are some discussions with JISC about some kind of support "to ensure the expertise and achievements of the AHDS are not lost to the community".
That leaves Performing Arts and Visual Arts; I can't even surmise what their future might be, since I don't know enough about their funding and local environment.
I appreciate that it's still early days, and no doubt crucial discussions are going on behind the scenes. But if any part of AHDS resources are in danger of loss, the resource owners need to consider plans to deal with those resources in the future. This will take some time, particularly since for more complex resources, it is clear that existing repositories are generally NOT yet adequate for purpose. I guess the picture itself will be complex; I can think of at least these categories:
- Some resources still exist outside of AHDS, and no action may be needed.
- Some resources will not be felt worth re-homing.
- Some resources can be re-homed in the time and funding institutions have available.
- Some resources should be re-homed, provided that time and funding are provided by some external source (this might be for developments on an institutional repository; it might be for work on the resource to fit a new non-AHDS environment).
- HOTBED (Handing on Tradition By Electronic Dissemination) pa-1028-1
- Lemba Archaeological Project arch-279-1
- Gateway to the Archives of Scottish Higher Education (GASHE) exec-1003-1
- Avant-Garde/Neo-Avant-Garde Bibliographic Research Database lll-2503-1
- Survey of Scottish Witchcraft, 1563-1736 hist-4667-1
- National Sample from the 1851 Census of Great Britain hist-1316-1
Are we OK? Is there more? Who knows! I think we need much better tools to tell what is "at risk", so that plans can start being made. Of course, this could be happening, maybe I'm just not "in the loop".
Will AHRC consider bids for funding transitional work? I certainly hope so, although I don't know how this might be done. JISC is (I believe) planning one last round of its Capital Programme. Will they include provisions to enhance repositories so as to take these more complex resources? I certainly hope so!
European e-Science Digital Repository Consultation
Philip Lord wrote to tell me that he and and Alison Macdonald are conducting a study for the European Commission “Towards a European e-Infrastructure for e-Science Digital Repositories”, (e-SciDR) – see www.e-SciDR.eu. This is a short study to summarize the situation regarding repositories in Europe and to propose policies for the Commission for repository development in Europe. As part of the study process the Commission is hosting a public consultation through a questionnaire... The letter inviting participation follows:
"Dear Sir, Dear Madam,[NB Safari on the Mac appears not to work with this questionnaire, but Firefox does.]
May I invite you, as key stakeholders, to contribute to the development of a knowledge society and digital infrastructure in Europe, by taking part in the Commission’s online public consultation on e-Science Digital Repositories which is available at http://ec.europa.eu/yourvoice/ipm/forms/dispatch?form=eSciDR."
"This consultation forms a key part of the e-SciDR study funded by the Commission into repositories holding digital data and publications for use in the sciences (in the widest sense encompassing disciplines from the humanities and social sciences to the life sciences).
Your answers will help identify needs, priorities and opportunities which the European Union, through the Commission, can help address and drive forward in the FP7 Capacity Programme and will provide an important input to developing future policy initiatives.
I would be grateful if you could respond to the consultation by no later than 30 July 2007.
All answers will be strictly confidential and anonymised.
If you would like to receive a summary of the consultation results, please tick the corresponding box on the questionnaire.
Best regards,
Mário Campolargo
Head of Unit GÉANT & e-Infrastructure"
Tuesday, 17 July 2007
A little more on very long term time series
I wrote earlier about a visit to Rothamsted Research, to talk with them about some of their very long term time series of agricultural research data (since 1843... digitised since 1991). Asking around, the prevailing wisdom seems to be to break the time series data when the nature of the data changes. Keep those time series, un-touched. Then build your overall time-series by a set of transformations on the original datasets, where the actual transformations are well-documented.
Of course if (as is perhaps the nature of such agricultural experiments) the nature of the data changes pretty well every year, then you have to keep a series of one-year data snapshots. And that sounds pretty much like what Rothamsted's batch "sheets" are doing.
Meanwhile, I'm continuing to look for something a bit more authoritative than "prevailing wisdom"! Someone has promised me a reference to some Norwegian economic or price time series kept since the 16th century, so I'm hopeful! Any hints from readers welcome...
Of course if (as is perhaps the nature of such agricultural experiments) the nature of the data changes pretty well every year, then you have to keep a series of one-year data snapshots. And that sounds pretty much like what Rothamsted's batch "sheets" are doing.
Meanwhile, I'm continuing to look for something a bit more authoritative than "prevailing wisdom"! Someone has promised me a reference to some Norwegian economic or price time series kept since the 16th century, so I'm hopeful! Any hints from readers welcome...
The National Archives and Microsoft join forces...
On 4 July, The National Archives and Microsoft announced a Memorandum of Understanding "ensuring preservation of the nation´s digital records from the past, present and into the future". Partly this relates to the standardisation of the Office Open XML format for Microsoft's products (I note that O'Reilly appear strongly in favour of this activity, seeing no conflict with the standardisation of Open Document Format). I had thought, through the rumour mill, that MS had provided (or was to provide) the specifications of obsolete file formats. This would fit with TNA's PRONOM file format registry, and their Seamless Flow programme. However, the key paragraph instead seems to be:
Should data people care? Apart from the documentation and other information that many will have stored in obsolete proprietary formats, it turns out that many store their data that way as well. The excellent JISC-funded StORe project did a survey of several disciplines, and got over 350 responses. Although one respondent said “God preserve us from idiots who archive data in proprietary commercial formats (Excel spreadsheets and MS-word documents)!”, 220 said they kept source data in spreadsheets and the same number kept them in word processed documents. These were the largest categories except images (228)!
Let's hope they are keeping them somewhere else as well...
"Today´s announcement sees Microsoft provide The National Archives with access to previous versions of Microsoft´s Windows operating systems and Office applications powered by Microsoft Virtual PC 2007. Virtual PC 2007 enables people to run multiple operating systems at the same time on the same computer. This allows The National Archives to configure any combination of Windows and Office from one PC, thereby allowing access to practically any document based on legacy Microsoft file formats. It is estimated that The National Archives will have to manage many terabytes of data in these formats."This sounds as if TNA must keep licensed versions of all older MS products, but then can run them under this emulation mode. This is a LOT better than nothing, but not as good as open access to the old file formats, with the ability to build additional tools that implies. But maybe there's more than is apparent in the press release?
Should data people care? Apart from the documentation and other information that many will have stored in obsolete proprietary formats, it turns out that many store their data that way as well. The excellent JISC-funded StORe project did a survey of several disciplines, and got over 350 responses. Although one respondent said “God preserve us from idiots who archive data in proprietary commercial formats (Excel spreadsheets and MS-word documents)!”, 220 said they kept source data in spreadsheets and the same number kept them in word processed documents. These were the largest categories except images (228)!
Let's hope they are keeping them somewhere else as well...
Monday, 16 July 2007
Open Data... Open Season?
Peter Murray Rust is an enthusiastic advocate of Open Data (the discussion runs right through his blog, this link is just to one of his articles that is close to the subject). I understand him to want to make science data openly accessible for scientific access and re-use. It sounds a pretty good thing! Are there significant downsides?
Mags McGinley recently posted in the DCC Blawg about the report "Building the Infrastructure for Data Access and Reuse in Collaborative Research" from the Australian OAK Law project. This report includes a substantial section (Chapter 4) on Current Practices and Attitudes to Data Sharing, which includes 31 examples, many from the genomics and related areas. Peter MR wants a very strong definition of Open Access (defined by Peter Suber as BBB, for Budapest, Bethesda and Berlin, which effectively requires no restrictions on reuse, even commercially). Although licences were often not clear, what could be inferred in these 31 cases generally would probably not fit the BBB definition.
However, buried in the middle of the report is a cautionary tale. Towards the end of chapter 4, there is a section on risks of open data in relation to patents, following on from experiences in the Human Genome and related projects.
The report goes on:
Are there other examples of these kinds of restrictions being imposed? Or of problems ensuing because they have not been imposed, and the data left open? (Note, I'm not at all advocating closed access!)
Mags McGinley recently posted in the DCC Blawg about the report "Building the Infrastructure for Data Access and Reuse in Collaborative Research" from the Australian OAK Law project. This report includes a substantial section (Chapter 4) on Current Practices and Attitudes to Data Sharing, which includes 31 examples, many from the genomics and related areas. Peter MR wants a very strong definition of Open Access (defined by Peter Suber as BBB, for Budapest, Bethesda and Berlin, which effectively requires no restrictions on reuse, even commercially). Although licences were often not clear, what could be inferred in these 31 cases generally would probably not fit the BBB definition.
However, buried in the middle of the report is a cautionary tale. Towards the end of chapter 4, there is a section on risks of open data in relation to patents, following on from experiences in the Human Genome and related projects.
"Claire Driscoll of the NIH describes the dilemma as follows:(The reference given is Claire T Driscoll, ‘NIH data and resource sharing, data release and intellectual property policies for genomics community resource projects’ Expert Opin. Ther. Patents (2005) 15(1), 4)
It would be theoretically possible for an unscrupulous company or entity to add on a trivial amount of information to the published…data and then attempt to secure ‘parasitic’ patent claims such that all others would be prohibited from using the original public data."
The report goes on:
"Consequently, subsequent research projects relied on licensing methods in an attempt to restrict the development of intellectual property in downstream discoveries based on the disclosed data, rather than simply releasing the data into the public domain."They then discuss the HapMap (International Haplotype) project, which attempted to make data available while restricting the possibilities for parasitic patenting.
"Individual genotypes were made available on the HapMap website, but anyone seeking to use the research data was first required to register via the website and enter into a click-wrap licence for the use of the data. The licence entered into, the International HapMap Project Public Access Licence, was explicitly modeled on the General Public Licence (GPL) used by open source software developers. A central term of the licence related to patents. It allowed users of the HapMap data to file patent applications on associations they uncovered between particular SNP data and disease or disease susceptibility, but the patent had to allow further use of the HapMap data. The licence specifically prohibited licensees from combining the HapMap data with their own in order to seek product patents..."Checking HapMap, the Project's Data Release Policy describes the process, but the link to the Click-Wrap agreement says that the data is now open. See also the NIH press release). There were obvious problems, in that the data could not be incorporated into more open databases. The turning point for them seems to be:
"...advances led the consortium to conclude that the patterns of human genetic variation can readily be determined clearly enough from the primary genotype data to constitute prior art. Thus, in the view of the consortium, derivation of haplotypes and 'haplotype tag SNPs' from HapMap data should be considered obvious and thus not patentable. Therefore, the original reasons for imposing the licensing requirement no longer exist and the requirement can be dropped."So, they don't say the threat does not exist from all such open data releases, but that it was mitigated in this case.
Are there other examples of these kinds of restrictions being imposed? Or of problems ensuing because they have not been imposed, and the data left open? (Note, I'm not at all advocating closed access!)
Subscribe to:
Posts (Atom)