Monday, 16 July 2007

Government responds to UK Science Funding petition

8,623 people signed the petition to the UK Government on the £68 million funding reduction to science. The Government has just published a response. After some phrases indicating that the science budget continues to rise, the key paragraph is:
"The Department of Trade and Industry had been facing a number of new and historic budgetary pressures which required action to keep within its budgets. Non-ring-fenced budgets had been reduced as far as possible, so the Department then had to consider its ringfenced budgets, including the Science Budget. It was decided to use part of the underspends in the Science budget that had been accumulated in previous years. This decision did not affect either the 2006-07 budget allocations, or the 2007-08 budget allocations , nor did it affect the commitments set out in the 10 Year Science and Investment Framework."
I suppose accumulated underspends might be in neither the 2006-07 budget nor in the 2007-08 budget, but the impact of the reduction has definitely been felt in the 2007-08 year!

The original petition asked the Government to "revise research funding via the DTI", or alternatively "I wish the Government to review its recent decision outlined [in the text]...". I guess that's a NO to the first but perhaps a yes to the second (a review doesn't necessarily mean a change).

Thursday, 12 July 2007

Very long term data

Rothamsted Research is an agricultural research organisation based near Harpenden in England. There are many interesting features of this organisation (only a few of which I know), including its “classical experiments". One of these, started in 1843, must surely be one of the longest-running experiments with resulting time-series data anywhere. I visited this week, spoke to Chris Rawlings and others handling the data for this “Broadbalk” experiment on wheat yields, and also a couple of scientists working on somewhat younger experiments collecting moths (1930s) and aphids (1960s) on a daily basis (using light traps and vacuum traps respectively).

Digital preservation theory tells us that digital data are at risk from various kinds of changes in the environment. Most often we focus on the media, or the risk of format incompatibility. OAIS rightly asks us to think about contextual metadata of various kinds, but also recognises that there is risk of semantic drift and/or semantic loss rendering once-clear resources incomprehensible (this is the business about the Designated Community and its Knowledge Base that I wrote about recently). The OAIS model seems to me (though some argue against this view) premised on a preservation or perhaps archival view: resources are ingested, preserved within the archive, and then disseminated at a later date. I’ve always had doubts about how easily this fits with a more continuous curation model, where the data are ingested, managed, preserved and disseminated simultaneously.

However, intuitively even in the curation situation I have described, if “long enough: time passes, many of the concerns of digital preservation will apply. And the Broadbalk wheat yield experiment since 1843 is certainly long enough. In that time there has been semantic drift (words mean different things), changes in units (imperial to metric), changes in plot size/granularity and plot labelling more than once, changes in what is measured (eg dropping wet hay weight while keeping dry yield) and how it is measured (the whole plot or a sample), changes in accuracy of measurement, and many changes in treatments. These are very serious changes that could significantly affect interpretation and analysis.

In the case of the aphids experiment, I saw a log book in which changes in interpretation etc were meticulously recorded. Unfortunately I don't have a copy of any pages from that book to describe in more detail the kinds of issues it raised. Although weekly aphid bulletins are made available (explicitly not "published"!), the book itself has, I believe, not been published. My guess is that this book forms a critical part of the provenance information by which the quality of the data can be judged.

Of course, the Broadbalk experiment has not been digital since 1843. I didn’t see the records, but the early ones would have been longhand in ledgers or notebooks, and later perhaps typed up. I get the impression it was quite systematic from the beginning, so converting it to an Oracle database in around 1991 was a feasible if major task. Now they are planning to convert it to a new database system and perhaps make the data more widely available, so they are asking questions about how they should better handle these changes.

By the way, in one respect the experiments continue to be decidedly non-digital. Since the beginning, they have been collecting soil samples from the plots, and now have over 10,000 of them! This makes a fascinating physical collection as a reference comparator for the digital records.

I should say that they have done a lot of hard thinking and good work about some of these issues. In particular, they have arranged all their input data into batches (known as “sheets”), which sounds pretty much self-describing in what is effectively a purpose-built data description language. This means they can roll back and roll forward their databases using different approaches. These sheets are all prepared using the assumptions of their time (remembering that some were prepared more than a hundred years after the data they contain was first recorded), but in theory there is enough information about these assumptions to make decisions. So once they have decided how to deal with some of these changes, they are able to do the best job possible of implementing it.

A Swedish forestry scientist once argued forcefully, in the discussion session to a presentation, that the original observations were sacrosanct, must be kept and should not be changed. Do what you like with the analysis, he suggested: re-run it with varied parameters, re-analyse with new models, (implicitly) make the sorts of changes suggested here. This approach would suggest starting a new time series dataset whenever there is a change of the sort we have been describing. That’s not exactly what has been happening with the Broadbalk experiment, but the sheets approach is if anything even finer grained.

I’m interested to collate experience from other long-running experiments that may have faced these and related issues. So far I have heard of the Harvard Forest project, part of the NSF Long Term Ecological Research (LTER) Network (thanks Raj Bose), and some very long-running birth cohort studies in the social sciences. Further suggestions on comparators and literature to look at would be welcome.

Open Access everyone? Or not?

I occasionally look at the OpenDOAR service, which list information about repositories, and check out those which claim to include data (the term they use is datasets, although it is possible that “other” might also be applicable!). Last time I looked there were 7 listed in the UK which claimed to collect datasets; 4 of these were institutional repositories, one is the Southampton eCrystals archive, the 6th is the National Digital Archive of Datasets (NDAD), which collects government datasets on behalf of The National Archive, and the 7th is Nature Precedings.

I used to think that was a creative piece of listing by NDAD, since OpenDOAR was created for institutional repositories, wasn’t it? But as time has gone on, I’ve begun to think NDAD are right; the old definitions of repositories were much too tight, and NDAD and many other data archives ought to be considered as repositories, and listed in these kind of resources. On that basis perhaps AHDS should have been listed (although they might not bother now) and, I thought, the UK Data Archive, part of the Economic and Social Data Service, should also be listed. After all, they are repositories, and open access is a good thing, isn’t it?

At the UKDA’s 40th birthday celebrations yesterday, it was clear that Kevin Schurer (Director of UKDA) doesn’t quite share that view. The view expressed certainly was not a total rejection of Open Access in favour of a commercial approach. But he was certainly arguing for some barriers between some data and the users. In particular, while access to UKDA data is free for certain categories of users (certainly UK academics and probably some others), ALL users are required to register. Kevin made a strong case for this being an advantage; registration means that he can report to his funders who his users are (including how many of them might be independent or Government researchers as well as academic ones), and can also monitor which datasets they are using. He knows, for example, that the user base has grown from 25 or so after the first 5 years to 45,000 (I don't remember the exact figure, but of that order) after the first 40 years.

The same registration mechanism (and associated authentication and authorisation mechanisms) also allow them to apply greater access controls to more sensitive datasets, including the possibility that the user may have to sign and observe special licences. It’s worth remembering that much of the data they hold is about people, and some is extremely sensitive information, which has been provided (under “informed consent”) for certain specific purposes.

In a very different environment the day before, I met some scientists with very long-term experiments collecting data on insects. For perhaps different reasons, they were happy to make their data available to collaborators, but not openly available on the web. Too many risks of mis-interpretation, which then requires extra effort to refute, was one reason. No doubt an extra paper as co-author from a collaboration was another motive (not unreasonable).

Are these approaches Open? Are they consistent with the OECD Principles and Guidelines for Access to Research Data that I wrote about earlier? Let’s remember that the Openness Principle was carefully worded:
“Openness means access on equal terms for the international research community at the lowest possible cost, preferably at no more than the marginal cost of dissemination. Open access to research data from public funding should be easy, timely, user-friendly and preferably Internet-based.”
So it looks as though the approach is reasonably consistent, provided access is on “equal terms for the international research community”, by that definition. I suspect there are other definitions (who said “the nice thing about standards is there are so many to choose from”?) by which these approaches would fail (eg the Berlin Declaration).

What the UKDA approach might do is make certain kinds of comparative work, including automated data mining, more difficult. There was a plea with respect to the former, from an Australian speaker, for more sophisticated international cross-archive access management (code: Shibboleth-enable). But I suspect with respect to the latter (preventing data mining) Kevin might argue “very right and proper too”!

Wednesday, 11 July 2007

UKDA 40th Birthday: back to basics?

There are not many digital data management organisations that can claim 40 years of continuous service. This week the UK Data Archive (UKDA, which holds mainly social science datasets) celebrates 40 years since its founding in 1967, with a party yesterday in the House of Commons followed by a small workshop in the UKDA’s fancy new quarters at the University of Essex. The DCC would like to congratulate the UKDA, Director Kevin Schurer and his 6 predecessor Directors, and the UKDA staff, on achieving this milestone.

[Pause to suppress pangs of envy at such longevity. Sigh!]

There were many very interesting aspects to this workshop, but here I will focus on one particular contribution, during the closing sessions, from Myron Gutmann. Myron is Director of ICPSR, a roughly equivalent organisation in the US, based at Michigan. Myron said he wanted to argue for retaining the basics. A data archive, he said, should do (at least) 5 things:
  • Appraise
  • Curate
  • Preserve
  • Train
  • Protect
This is not quite the language of OAIS, but perhaps closer to the language of archives. Appraise because everything an archive does is expensive, so it has a responsibility to select high quality resources appropriate to its mission (and as OAIS might suggest, its Designated Community). Curate because (apparently) even the best resources tend to need extensive work to make the usable (this work includes making the contextual and other metadata about the resource usable by non-insiders; we might perhaps describe this perhaps as preservation metadata, maybe as part of its representation information). Preserve for obvious reasons, to make resources available in the long term (although some people at the workshop seemed to be using the word in a special sense: institutional repositories are not about preservation, said someone, although I may have mis-quoted). Train, because these resources are specialist, and often need significant work and extra knowledge on the part of users. And protect, both as protection of the intellectual property in the resource, but more importantly protecting the subjects of the data, whose sensitive data must be made available only under appropriate conditions.

It’s a pretty good summary of the job of an archive, give or take a verb or two. Myron added something like “serve new user bases” and “innovate” in various ways, and it’s hard to argue with.

There were some sober reflections on sustainability at the workshop, partly relating to the difficult funding position for long-term longitudinal surveys, but partly reflecting the recent AHRC decisions. Paraphrasing Kevin Schurer, we never could or should take things for granted, but now we must be doubly persuasive, even if the UKDA's main funder ESRC regards ESDS (of which UKDA is a key part) as “a jewel in its crown”.

Friday, 6 July 2007

On blog authorship and (un)certainty

There's a difference between a blog and an article, and it seems to me it's about certainty. Why should I write this blog? It could be to record trivial events, it could be self-aggrandisement, but I think it's about dealing with uncertainty. If I were fully convinced about the detail, I guess I would write an article and submit it to a journal. But generally I'm speculating a bit; trying to focus my mind by writing something out as clearly as I can for an unknown audience (an audience that has the power to answer back). There's also a nice comment in David Rosenthal's blog: "Taking small, measurable steps quickly is vastly more productive than taking large steps slowly, especially when the value of the large step takes even longer to become evident".

I know I'm writing things that colleagues may not agree with. Sometimes, I expect the process of writing to move me closer to their ideas. Sometimes I might hope they move closer to my ideas. Sometimes we may have to learn to live with our differences. The point for me is to generate some strange kind of conversation, not just with close colleagues but to get some feedback from colleagues interested in data curation across the world.

If I can work out what is causing confusion, articulate it and resolve it, perhaps I can stop being "confused of Kenilworth"!

Representation Information: what is it and why is it important?

Representation Information is a key and often misunderstood concept. To understand it, we need to look at some definitions. First of all, OAIS (CCSDS 2002) defines data thus:
“Data: A reinterpretable representation of information in a formalized manner suitable for communication, interpretation, or processing.”
Second, we have Information:

“Information: Any type of knowledge that can be exchanged. In an exchange, it is represented by data. An example is a string of bits (the data) accompanied by a description of how to interpret a string of bits as numbers representing temperature observations measured in degrees Celsius (the representation information).”
Then we have Representation Information (sometimes abbreviated as RI):
“Representation Information: The information that maps a Data Object into more meaningful concepts. An example is the ASCII definition that describes how a sequence of bits (i.e., a Data Object) is mapped into a symbol.”
As an example, we have this paragraph:
"Information is defined as any type of knowledge that can be exchanged, and this information is always expressed (i.e., represented) by some type of data. For example, the information in a hardcopy book is typically expressed by the observable characters (the data) which, when they are combined with a knowledge of the language used (the Knowledge Base), are converted to more meaningful information. If the recipient does not already include English in its Knowledge Base, then the English text (the data) needs to be accompanied by English dictionary and grammar information (i.e., Representation Information) in a form that is understandable using the recipient’s Knowledge Base.”
The summary is that “Data interpreted using its Representation Information yields Information”.

Now we have a key complication:
“Since a key purpose of an OAIS is to preserve information for a Designated Community, the OAIS must understand the Knowledge Base of its Designated Community to understand the minimum Representation Information that must be maintained. The OAIS should then make a decision between maintaining the minimum Representation Information needed for its Designated Community, or maintaining a larger amount of Representation Information that may allow understanding by a larger Consumer community with a less specialized Knowledge Base. Over time, evolution of the Designated Community’s Knowledge Base may require updates to the Representation Information to ensure continued understanding.“
So now we need another couple of definitions:
“Designated Community: An identified group of potential Consumers who should be able to understand a particular set of information. The Designated Community may be composed of multiple user communities.”
“Knowledge Base: A set of information, incorporated by a person or system, that allows that person or system to understand received information.”
So there are several interesting things here. First is that this obviously enshrines a particular understanding of information; one I couldn’t find in Wikipedia when I last looked (here is the article at that time; may be it will be there next time!). Floridi suggests there is no commonly accepted definition of information, and that it is polysemantic, and particularly contrasts information in Shannon’s Mathematical Theory of Communication with the “Standard Definition of Information” (Floridi, 2005). If I understand it rightly, the latter refers to factual information (with some controversy on whether it need be true), but not necessarily to instructional information (“how”).

Secondly, the introduction of the Designated Community and its Knowledge Base may be both helpful and problematic. It may be helpful because it can reduce the amount of Representation Information needed to interpret data (or even eliminate it completely). This is if the Designated Community is defined as having a Knowledge Base that allows it to understand the data, then nothing more is required. This is obviously never entirely true, and in practice even with a Designated Community that is quite strongly familiar with the data, we will expect to need some RI, perhaps to identify the particular meaning of some variables, etc.

The problematic nature arises because we now have two external concepts, the Designated Community and its Knowledge Base that influence what we must create, and which will change and must be monitored. I’ve heard the words “precise definition” used in the context of these two terms, but I am sceptical anyone can define either precisely (although the LOCKSS Statement of Conformance with OAIS has a minimalistic go; it's the only public one I could find, but I would love to see more). My colleague David Giaretta suggests that his huge project CASPAR aims to produce better definitions.

In fact, although they may be useful ideas, both the Designated Community and its Knowledge Base seem to be quite worrying terms. The best we can say is that “chemists” (for example) understand “”chemical concepts”, and that the latter have proved pretty stable at least in basic forms. But the community of chemists turns out to include a myriad of sub-disciplines, with their own subtleties of terminology, and not surprisingly introducing new concepts and abandoning old ones all the time. If we have some chemical data in our repository, we have to watch out for these concepts going from current through obsolescent, obsolete to arcane, and in theory we have to add RI at each change, to make up for the increasing gap in understanding.

The third interesting feature is that these definitions say nothing about files or file formats at all, yet “format registries” are the most common response to meeting the need for RI. TNA’s PRONOM, and the Harvard/OCLC Global Digital Format Registry (GDFR) are the two best-known examples.

Clearly files and file formats play a critical role in digital preservation. Sometimes I think this has occurred because of the roots of much of digital preservation (although not OAIS) lie in the library and cultural heritage communities, dominated as they are by complex proprietary file formats like Microsoft Word. In science, formats are probably much simpler overall, but other aspects may be more critical to “understand” (ie use in a computation) the data.

The best example I know to illustrate the difference between file format information and RI is to imagine a social science survey dataset encoded with SPSS. We may have all the capabilities required to interpret SPSS files, but still not be able to make sense of the dataset if we do not know the meaning of the variables, or do not have access to the original questionnaires. Both the latter would qualify as RI. Database schemas may provide another example of RI.

Have I shown why or how RI is a useful concept in digital curation? I'm not sure, but at least there's a start. Representation Information, as David Giaretta sometimes says, is useful for interpreting unfamiliar data!

In later posts, I’m going to try to include some specific examples of RI that relates to science data. I also intend to try to justify more strongly the role of RI in curation rather than preservation, ie through life rather than just at the end of it!


* CCSDS (2002) Reference Model for an Open Archival Information System (OAIS). IN CCSDS (Ed.), NASA.

* FLORIDI, L. (2005) Is Semantic Information Meaningful Data? Philosophy and Phenomenological Research, 70, 351-370. http://www.ingentaconnect.com/content/ips/ppr/2005/00000070/00000002/art00004

Authenticity across migrations

I discovered a few days ago that I have 4 digital objects that are (I believe, but am not certain) in some strong senses “the same” (in their information content), but which are also completely different (in their bits). These objects are the result of a chain of “exports” and “imports”, and “save as…” operations, prompted partly by a change of technology (from a Windows PC running Mind Manager to a Macintosh running NovaMind), and partly from a need to make the content of the object more accessible to colleagues who do not use either software package. In case you're interested, these files are now accessible at the URLs below; I have shown the original date for each file:

15/10/2004: http://www.dcc.ac.uk/docs/blog/GWG%20vision%20and%20action.mmp
09/08/2005: http://www.dcc.ac.uk/docs/blog/GWG%20vision%20&%20action.xml
06/10/2005: http://www.dcc.ac.uk/docs/blog/GWG%20vision%20and%20action.pdf
10/01/2006: http://www.dcc.ac.uk/docs/blog/GWG%20vision%20&%20action.nmind

I’m very interested in the question of authenticity across migrations. Whilst migrations for preservation purposes are perhaps likely to be needed much less often than we once thought, they are nevertheless inevitable given long enough time. How can we plausibly assert that these objects represent “the same thing”?

In this case, each of these objects represents a simple document, that is they can each be transformed into a 2-D representation. Three of the objects have other capabilities: with the appropriate software they can be edited, by adding or removing elements, and by moving elements in relation to one another. This is true to some extent of the fourth object (a PDF, which for most people can be viewed to see a representation of the object), but in a different and much more restricted way. However, this editability, while vital for some kinds of re-use, is not essential for conveying the information essence of the objects.

These documents happen to be mind maps. The PDF should help some of those unfamiliar with this type of document. Mind mapping is a particular approach to organising ideas that originated from Tony Buzan (Buzan 1974). I find the approach particularly useful and productive; however I suspect most of my colleagues groan inwardly when yet another mind map is produced for them to consider. They are a little lightweight as communications devices (and so may have rather limited value as records), yet I persist in thinking they may have value for others. An interesting feature here, however, is that there is no real standardisation in this market. Mind Manager (from MindJet) may perhaps have the market share to force some degree of standardisation (hence the ability of NovaMind and at least one of the Open Source versions that I am aware of , ie FreeMind, to import from Mind Manager XML files). This fractured market led to my having these 4 versions.

I realise that this blog post runs a real risk of being dismissed because of a naïve and undefined approach to “sameness”, and that this idea of sameness, is likely to be linked to the “significant properties” of the object; further that these significant properties will differ for different users. I’ll try to re-visit these ideas later (for now, note a report by Hedstrom and Lee, and a recent ITT from JISC).

One approach to asserting authenticity in this case might be to print out or view each document, and use a visual comparison to assert sameness. It’s worth noting though that I cannot use this technique easily at the moment across the full set, because I no longer have access to the application software that will render the first of the objects (although I believe the test could be done; the object is not obsolete). Furthermore, although the comparison is feasible for such a small number of objects, it is not scalable to the large numbers of objects that one would expect to find in a typical digital repository environment.

One other technique could perhaps be to record provenance data on the processes involved in the migration. In this case the object was created with an old version of Mind Manager, and then exported as a Mind Manager XML file once the conversion to Mac had occurred, and it appeared that no Mac version of Mind Manager was likely (although two years later one exists). This XML version was created explicitly as a migration stage, because it could be read an imported by the Mac application of choice, the software NovaMind. At some point, a PDF version was created to allow the object to be emailed to colleagues who did not have either application software. Shorn of dates and evidence, that is the provenance chain for these 4 objects… except I have a niggling concern that I might have changed the first Novamind version after importing it. A true provenance chain would record ALL the changes to the object; in this case, I have found a way of checking, but I’ll leave it as an exercise for the reader to work out whether it was changed or not! The problem is that this provenance chain only tells us what happened at a gross level (what large scale operations were applied), not whether the migrations were successful, nor what artefacts were introduced or features lost.

Perhaps there is a way, maybe linked to particular families or genres of document types, of computing properties from documents that we would strongly wish to see as stable across migrations. In this case, the text labels on the arms of the mind maps might be one such desirable invariant; quite a strong one, but not perfect (for example, the text labels may show as invariant but might have become detached from their branches, destroying the meaning in their relationships, while more adventurous mind mappers use much more in the way of imagery and other techniques to convey part of their message).

However, if we are to plausibly assert authenticity across migrations, we will have to identify some such invariants. They may be extremely important parts of the preservation metadata, just as checksums and signatures are for checking the authenticity of objects that we believe have not been changed. Any suggestions on candidate invariants for some object classes?

* BUZAN, T. (1974) Use Your Head, BBC.