Monday, 10 December 2007

Two fundamentally different views on data curation

A few months ago, we interviewed two scientists for two quite different posts in the DCC. Both were from a genomics background, and I was very struck by how strongly held, but from my point of view how narrow, was their view of curation. As a result, I’m beginning to realise that there are two fundamentally different approaches to data curation

The Wikipedia definition of biocurator was quoted by Judith Blake, the keynote speaker [slides] at the 2nd International BioCuration Meeting in San Jose:
“A biocurator is a professional scientist who collects, annotates, and validates information that is disseminated by biological and model organism databases. The role of a biocurator encompasses quality control of primary biological research data intended for publication, extracting and organizing data from original scientific literature, and describing the data with standard annotation protocols and vocabularies that enable powerful queries and biological database inter-operability. Biocurators communicate with researchers to ensure the accuracy of curated information and to foster data exchanges with research laboratories.

Biocurators (also called scientific curators, data curators or annotators) have been recognized as the ‘museum catalogers of the Internet age’ [Bourne & McEntyre]”
In essence, this suggests that curation (let’s call it biocuration for now) is the construction of authoritative annotations linking significant objects (eg genes) with evidence about them in the literature and elsewhere. And this is indeed the sort of thing you see in genomics databases.

In other parts of science, I believe curation is not so much about constructing annotation on objects, but rather about caring for arbitrary datasets (adding descriptions, information about how to use them, transferring them to longer term homes, clarifying conditions of use and provenance information, but yes including annotations, etc), so that they can be used and re-used, by the originators or others, now or in the future. This approach incorporates elements of digital preservation, but is not solely defined by “long term” in the way that digital preservation tends to be. Clearly we will find cases where we need to curate datasets which themselves result from biocuration!

The problem is that both terms are well-embedded in their respective communities. I don’t think either of my two interviewees had any idea how much broader our concept was than theirs. We need to watch out for the inevitable misunderstandings that can arrive from this clash of meanings! However, both communities are likely to continue to use the terms in their own ways (although perhaps this term biocurator, relatively new to me, is an effort to disambiguate).

Sunday, 9 December 2007

Posting gap… and intriguing article on “borrowed data”

Apologies to those of you who were interested in this blog; there has been only one posting since the end of August. Shortly after that I went off-air after an accident, and since then I have been recovering. Although not yet officially back at work, I hope I have now reached the point where I can start making some postings again.

To alleviate potential boredom, a friend gave me some back issue of New Scientist to read. The first one I opened was the issue for 20 January, 2007. On page 14, the first paragraph of an article titled “Loner stakes claim to gravity prize” jumped out at me. It read:
“A lone researcher working with borrowed data may have pipped a $700 million NASA mission to be the first to measure an obscure subtlety of Einstein’s general theory of relativity.”
The thrust of the New Scientist piece is that this “lone researcher” had re-analysed another scientist’s analysis of the orbit of a NASA satellite of Mars, and claimed to have found evidence of the Lense-Thirring effect, a twisting of space-time near a large rotating mass, ahead of NASA’s own expensive mission.

Borrowed data? This had to be worth following up! The relevant article is “Testing frame-dragging with the Mars Global Surveyor spacecraft in the gravitational field of Mars” by Lorenzo Iorio, available from http://arxiv.org/abs/gr-qc/0701042. If you check that reference you may note that 3 versions of the article had been deposited by the cover date of the New Scientist piece, but that he is now up to version 10, deposited on 14 May 2007. Checking the “cited by” citations in SLAC-SPIRES HEP shows that this article has been very controversial, with many comments and replies to comments, apparently contributing to the development of the article (although the authors of these comments don’t get any acknowledgments as far as I could see). Nevertheless, it does look like a nice example of the value of an open approach to science. I don’t know how much the final article is now accepted, although it does seem now to form part of a book chapter (L. Iorio (ed.) The Measurement of Gravitomagnetism: A Challenging Enterprise, Chap. 12.2, NOVA Publishers, Hauppauge (NY), 2007. ISBN: 1-60021-002-3).

What about the borrowed data? The Acknowledgments section has: “I gratefully thank A. Konopliv, NASA Jet Propulsion Laboratory (JPL), for having kindly provided me with the entire MGS data set”. He does not give an explicit data citation. It appears however, that Iorio is not referring to the publicly accessible science results accessible from JPL. In section 2, he writes:
“In [24] six years of MGS Doppler and range tracking data and three years of Mars Odyssey Doppler and range tracking data were analyzed in order to obtain information about several features of the Mars gravity field summarized in the global solution MGS95J. As a by-product, also the orbit of MGS was determined with great accuracy.”

Reference 24 is “Konopliv A S et al 2006 Icarus 182 23”, which appears to be “A global solution for the Mars static and seasonal gravity, Mars orientation, Phobos and Deimos masses, and Mars ephemeris”, by Alex S. Konopliv, Charles F. Yoder, E. Myles Standish, Dah-Ning Yuan and William L. Sjogren., in Icarus Volume 182, Issue 1, May 2006, Pages 23-50. In turn, this paper acknowledges “Dick Simpson and Boris Semenov provided much of the MGS and Odyssey data used in this paper mostly through the PDS archive”, and this time there is a relevant data citation: “Semenov, B.V., Acton Jr., C.H., Elson, L.S., 2004a. MGS MARS SPICE KERNELS V1.0, MGS-M-SPICE-6-V1.0. NASA Planetary Data System”. However this presumably represents the source data from which Konopliv et al did their calculations.

(BTW I have yet to find a formal place where JPL PDS define their requirements for data citations, but there are a couple of examples in http://pds.jpl.nasa.gov/documents/pag/MERDSC.doc, and the citation above sticks to that guideline.)

It looks like the “borrowed data” is in fact provided on a colleague-to-colleague basis. It is clear from some of the replies that there was correspondence between Iorio and Konopliv on some matters of interpretation.

Nevertheless, this seems like a couple of useful examples of science being discovered from the analysis of original and derived datasets established for another purpose, and hence of relevance to data curation. And who knows, maybe the $700 million Gravity Probe mission launched by NASA will turn out not to have been necessary after all?

BTW I tried to follow this story through New Scientist’s online service, to which my University has a subscription. I was asked for my ATHENS credentials, which I provided, but was then thrown out on an IP address check, despite using a VPN. How not to get good use of your magazine!

Wednesday, 31 October 2007

Digital Curation Centre/London eScience Centre Collaborative Workshop

A post from Graham Pryor (DCC eScience Liaison)

The format of collaborative workshops having been set before I picked up the DCC’s eScience Liaison mantle in July this year, I was of course eager to match the success of previous events. My first, with the London eScience Centre (LeSC), was scheduled for 16th October, and in the spirit of exchanging knowledge and expertise, a clutch of presentations and demonstrations by DCC staff was assembled that could be expected to showcase the range of activities and services in which we are engaged. But how would they meld with contributions from the LeSC and how appropriate would be the ‘collaboration’?

I had no cause for concern. The LeSC, now transformed in name and context as the Imperial College Internet Centre, had welcomed this event, producing a spectrum of topics and speakers ranging from an exposition of the College’s ICT and Library collaboration on digital repositories to Henry Rzepa’s demonstration of a multifaceted digital repository for chemistry and Omer Casher’s ‘Semantic Eye’. Opening the workshop, the Internet Centre’s Director, John Darlington, described its new remit as embracing the gamut of Internet-enabled research and development, a significantly broader mission than that had been suggested by the superseded title of London eScience. In fact, the agenda had been evolving up to the last minute, emerging as a very full and lengthy event that overturned the format of previous events by displacing the concluding plenary discussion, but with questions allowed for after each presentation we did perhaps enjoy a greater sense of immediacy and relevance in the lively debates that ensued. (Presentations from the workshop are to be made available at the Internet Centre’s Web pages)

For the DCC, these discussions proved to have particular relevance to the compelling subject of the project’s sustainability, as we found ourselves dealing with questions regarding the range and level of service provision that could be expected from the DCC – e.g. ‘when a project is at start-up can the investigators rely on the provision of an advisory service covering legal and other core data curation issues?’ – which very much reflects the shift in emphasis of DCC Phase 2.

Closing the workshop, John Darlington judged the event to have been a success, having brought together not only the Internet Centre and the DCC, but also colleagues from within Imperial College who might otherwise not have an opportunity to discuss areas of shared interest. His perception was that the workshop had fully supported the rationale of the Internet Centre and he was enthusiastic to maintain the impetus for the exchange of knowledge that it had provided. Upon reflection, that is another box to be ticked on the account of DCC service provision, but the larger message of how should we focus our portfolio for future services and support is one that I took back with me.

Friday, 31 August 2007

Archiving Service

The other day I had an interesting discussion with a group of IT facilities management guys here at Edinburgh, who have been asked to write the requirements for a new version of their Archive Service, capable of handling modern requirements for huge amounts of data. It was an interesting discussion, and I hope to remain involved in what they are planning (which is a long way from getting funded, I guess). They had some good ideas, thinking perhaps derived in part from software management systems, of checking data in and out, managing versions, etc.

We spoke a bit about OAIS, and about the issues of making data understandable for the long term. They were somewhat reluctant to go there, not surprisingly; general IT attitude to "archiving" in the past has tended to be something like: "You give me your blob of data, and I'll give you back an identical blob of data at some future time". IT people are pretty comfortable about that approach (although I also mentioned some emerging issues such as those mentioned by David Rosenthal in his blog post about keeping very large data for very long times).

We also discussed the need to make the data accessible from outside the University, to allow it to be linked from publications. This too was a bit outside their previous remit.

It struck me afterwards how much the conversation was affected by the IT services point of view (I'm not trying to denigrate these individuals at all; as far as I can see they took it all on board). I have talked with a lot of people about "digital preservation" or "repositories"; in most of those conversations, the issues about managing the bits themselves are assumed to be pretty much a solved problem, and the conversation takes different forms, about metadata, about formats, about representation information, about emulation versus migration, and so on. The participants tend to be linked to the library, or to some subject area, but rarely from IT.

I wondered if it was possible to imagine the IT guys providing a secure substrate on which a repository service sits, with independence between the two; that would allow everyone to stay in their comfort zone. I don't think OAIS would quite work in those terms, and I think there would be dependencies both ways, although I'll have to think more about this. The only counter example I have heard of is a rumour of Fedora providing high level services based on SRB as a secure storage substrate, although I don't have a reference.

What was also interesting is how few examples we could think of in Universities, of large scale, long term archiving systems for records and data, leaving aside the publication-oriented repository movement.

Can any readers give us pointers?

Wednesday, 29 August 2007

IJDC again

At the end of July I reported on the second issue of the International Journal of Digital Curation (IJDC), and asked some questions:
"We are aware, by the way, that there is a slight problem with our journal in a presentational sense. Take the article by Graham Pryor, for instance: it contains various representations of survey results presented as bar charts, etc in a PDF file (and we know what some people think about PDF and hamburgers). Unfortunately, the data underlying these charts are not accessible!

"For various reasons, the platform we are using is an early version of the OJS system from the Public Knowledge Project. It's pretty clunky and limiting, and does tend to restrict what we can do. Now that release 2 is out of the way, we will be experimenting with later versions, with an aim to including supplementary data (attached? External?) or embedded data (RDFa? Microformats?) in the future. Our aim is to practice what we may preach, but we aren't there yet."

I didn't get any responses, but over on the eFoundations blog, Andy Powell was taking us to task for only offering PDF:
"Odd though, for a journal that is only ever (as far as I know) intended to be published online, to offer the articles using PDF rather than HTML. Doing so prevents any use of lightweight 'semantic' markup within the articles, such as microformats, and tends to make re-use of the content less easy."
His blog is more widely read than this one, and he attracted 11 comments! The gist of them was that PDF plus HTML (or preferably XML) was the minimum that we should be offering. For example, Chris Leonard [update, not Tom Wilson! See end comments] wrote:
"People like to read printed-out pdfs (over 90% of accesses to the fulltext are of the pdf version) - but machines like to read marked-up text. We also make the xml versions availble for precisely this purpose."
Cornelius Puschmann [update, not Peter Sefton] wrote:
"Yeah, but if you really want semantic markup why not do it right and use XML? The problematic thing with OJS (at least to some extent) is/was that XML article versions are not the basis for the "derived" PDF and HTML, which deal almost purely with visuals. XML is true semantic markup and therefore the best way to store articles in the long term (who knows what formats we'll have 20 years from now?). HTML can clearly never fill that role - it's not its job either. From what I've heard OJS will implement XML (and through it neat things such as OpenOffice editing of articles while they're in the workflow) via Lemon8 in the future."
Bruce D'Arcus [update, not Jeff] says:
"As an academic, I prefer the XHTML + PDF option myself. There are times I just want to quickly view an article in a browser without the hassle of PDF. There are other times I want to print it and read it "on the train."

"With new developments like microformats and RDFa, I'd really like to see a time soon where I can even copy-and-paste content from HTML articles into my manuscripts and have the citation metadata travel with it."
Jeff [update, not Cornelius Puschmann] wrote:
"I was just checking through some OJS-based journals and noticed that several of them are only in PDF. Hmmm, but a few are in HTML and PDF. It has been a couple of years since I've examined OJS but it seems that OJS provides the tools to generate both HTML and PDF, no? Ironically, I was going to do a quick check of the OJS documentation but found that it's mostly only in PDF!

"I suspect if a journal decides not to provide HTML then it has some perceived limitations with HTML. Often, for scholarly journals, that revolves around the lack of pagination. I noticed one OJS-based journal using paragraph numbering but some editors just don't like that and insist on page numbers for citations. Hence, I would be that's why they chose PDF only."
I think in this case we used only PDF because that was all our (old) version of the OJS platform allowed. I certainly wanted HTML as well. As I said before, we're looking into that, and hope to move to a newer version of the platform soon. I'm not sure it has been an issue, but I believe HTML can be tricky for some kinds of articles (Maths used to be a real difficulty, but maybe they've fixed that now).

I think my preference is for XHTML plus PDF, with the authoritative source article in XML. I guess the workflow should be author-source -> XML -> XHTML plus PDF, where author-source is most likely to be MS Word or LaTeX... Perhaps in the NLM DTD (that seems to be the one people are converging towards, and it's the one adopted by a couple of long term archiving platforms)?

But I'm STILL looking for more concrete ideas on how we should co-present data with our articles!

[Update: Peter Sefton pointed out to me in a comment that I had wrongly attributed a quote to him (and by extension, to everyone); the names being below rather than above the comments in Andy's article. My apologies for such a basic error, which also explains why I had such difficulty finding the blog that Peter's actual comment mentions; I was looking in someone else's blog! I have corrected the names above.

In fact Peter's blog entry is very interesting; he mentions the ICE-RS project, which aims to provide a workflow that will generate both PDF and HTML, and also bemoans how inhospitable most repository software is to HTML. He writes:
"It would help for the Open Access community and repository software publishers to help drive the adoption of HTML by making OA repositories first-class web citizens. Why isn't it easy to put HTML into Eprints, DSpace, VITAL and Fez?

"To do our bit, we're planning to integrate ICE with Eprints, DSpace and Fedora later this year building on the outcomes from the SWORD project – when that's done I'll update my papers in the USQ repository, over the Atom Publishing Protocol interface that SWORD is developing."
So thanks again Peter for bringing this basic error to my attention, apologies to you and others I originally mis-quoted, and I look forward to the results of your efforts! End Update]

OAIS review: what's happening?

In June 2006, there was an announcement:
In compliance with ISO and CCSDS procedures, a standard must be reviewed every five years and a determination made to reaffirm, modify, or withdraw the existing standard. The “Reference Model for an Open Archival Information System (OAIS)” standard was approved as CCSDS 650.0-B-1 in January 2002 and was approved as ISO standard 14721 in 2003. While the standard can be reaffirmed given its wide usage, it may also be appropriate to begin a revision process. Our view is that any revision must remain backward compatible with regard to major terminology and concepts. Further, we do not plan to expand the general level of detail. A particular interest is to reduce ambiguities and to fill in any missing or weak concepts. To this end, a comment period has been established.
Comments were required by 30 October 2006. The Digital Preservation Coalition and the Digital Curation Centre ran a joint workshop on 13 October in Edinburgh, and as a result submitted joint comments. Some of these comments were minor but some were quite significant. There are 14 general recommendations, and many detailed updates and clarifications were suggested.

Supporters of OAIS (and I am one) often make a great play about how open their process was. In that spirit, I have tried to find out what is happening, and to take part. I understood there was to be an open process, with a wiki and telecons, to decide in the first place whether the standard needs revision, and if so to revise it. Despite numerous attempts, I cannot find out what the current state is, or where or how this is taking place. Does anyone know?

This is a very important standard, and it needs revision to make it more useful in today's environment. We need to make it as good as we can.

Wednesday, 22 August 2007

Proportion of research output as data or publication?

My colleague Graham Pryor asks in an email:
"Chris - I have been looking for evidence of the proportion of UK research output that can be categorised as scholarly publications and that which is in the form of data. I have found nothing. It is quite possible that no-one has ever tried to work this out. However, on the off-chance, is this a figure (even an estimate) that you might have come across?"
I think this is a great question, even an "Emperor's New Clothes" question. "Everyone knows" that the data proportion has been increasing, but I know of no estimates of the proportions.

Does anyone else know of such estimates? If not, does anyone have any idea how to set about making such an estimate? A very clear definition of data would be needed, to include externally usable data perhaps, rather than raw telemetry...

(Apologies for the gap in posting, I have been sequestered on a Greek island for a couple of weeks, and not thinking of you at all!)