Showing posts with label Digital Curation. Show all posts
Showing posts with label Digital Curation. Show all posts

Tuesday, 5 January 2010

Digital Curation Centre User Survey 2009: Highlights

My colleague Angus Whyte has provided the following brief summary of two surveys carried out in Phases 1 and 2 of the Digital Curation Centre, in 2006 and 2009 respectively, as part of our evaluations. In retrospect, we might have done better revising the questions for the second survey rather more than we did; nevertheless I thought it worth while sharing this with you.

Angus writes:

In 2009 DCC users were surveyed, repeating a similar survey carried out in 2006. In the highlights below we draw conclusions both from the more recent results and also changes over the 3 year period. Both surveys were publicised on the DCC website and via several mailing lists, principally the DCC-Associates and (in 2009) the JISC sponsored Research-Dataman list.

Our conclusions take into account that the online questionnaire was self-completed by a self-selected group of respondents (75 in 2009 and 125 in 2006). DCC Associates (640 approx.) provided the bulk of the responses[1]. The results indicated broad patterns, relatively wide differences and consistent responses over the two surveys, even though these are not taken to be statistically representative.

Highlights

In both surveys around 90% of respondents are familiar with the term ‘digital curation’ and regard it as a critical issue within their project or unit. The DCC is consistently given as the main source of information on curation issues by around 70% of respondents, with “on the job challenges/ research” second at around 60%.

Between the two surveys there is a large jump (from 13% to 32%) in the number of respondents indicating that DCC has been “very effective” in raising awareness about digital curation, and those believing it to be “slightly effective” has correspondingly fallen from 53% to 31%.

Of a list of DCC resources, five are identified as “most helpful” by at least 1 in 5 of the 2009 survey respondents, these being (in descending order) the DCC website, Briefing Papers (of various sorts), the DCC Curation Lifecycle Model, Case Studies, and the Digital Curation Manual.

Respondents universally associate digital curation with “ensuring the long-term accessibility and re-usability of digital information”, and large majorities (around 90%) also relate it to “performing archiving activities on digital information such as selection, appraisal and retention” and “ensuring the authenticity, integrity and provenance of digital information are maintained over time”. Rather lower but still significant numbers (around 60%) associate digital curation with “managing digital information from its point of creation” and “managing risks to digital information” – although many more highlight the latter in 2009 (up to 84% from 61%).

Curation or preservation addresses risks to the respondents’ organisations with “loss of organisational memory” consistently topping their list (identified by around 75% of respondents) and “business risks” second, identified by just under half, again across both surveys.

More than two thirds indicate that their main reasons for curating and preserving digital information are its educational/research or historical value; in both years a minority cites other reasons. Similarly, the main obstacles are indicated as financial or staff resources, with around half also indicating lack of awareness or appropriate policies.

For around 40% of respondents, management and preservation of digital information has an indefinite timescale. For a further 15% or so it is “beyond the life of the project/organisation”, and similar numbers indicate these are tasks “for the life of the project/organisation”.

The 2009 survey respondents are no strangers to the ‘data deluge’, most dealing with at least 100Gb and some (7%) more than 100Tb. Overall 79% expect this to increase in the next two years, surprisingly 3% do not, while 7% do not know. Most need to manage a mixture of open and proprietary formats, and report a wide variety of formats in use, predominantly common office applications, PDF documents and multimedia formats. Curation and preservation challenges are most frequently identified with obsolete proprietary formats. Image, video, and geospatial data are also often identified as challenges, as are web sites combining these.

Respondents were also asked in 2009 about re-use, and around a third indicate that research data is re-used internally, with similar numbers offering data generated by their project/unit for re-use by others, or re-using external data.

Access issues facing research projects/units are identified in both surveys and along similar lines; intellectual property rights (e.g. copyright) is the most frequently cited issue, followed by “privacy or ethical issues”, however “embargo on research findings” is least prevalent, identified by only a fifth of respondents.

Asked about funding for curation and preservation, responses show no clear picture. Around half of 2009 respondents indicate funding is “accounted for in project or institutional budget”. A large minority have no explicit funding for curation and preservation, and where resources are available these are pooled from other funded areas (e.g. IT budget for project or organisation) or research grants. Spending on curation/preservation is less than £50,000 (for around half of those respondents who were aware of this). Around half are unsure whether spending will increase or decrease, with the remainder being evenly split.

Detailed questions and response data are available on request.

Angus Whyte, Digital Curation Centre

[1] The DCC Associates membership list includes UK data organisations, leading data curators, overseas and supranational standards agencies, and industrial/business communities. Currently research data creators are under-represented (information from registration details).


Tuesday, 8 December 2009

Leadership opportunities

Those interested in leadership in Digital and Data Curation should keep an eye on the relevant UK press and lists over the next week or so for anything of interest...

Monday, 10 August 2009

Forgetting to remember

After a Sunday Times article prompted yesterday's piece of whimsy, a Tweet from my standard Twitter search ( (digital OR data) AND (preservation OR curation), since you ask) produced an interesting article by Chris O'Brien, a columnist for MercuryNews.com: "Time to clean up your digital closet". He goes quite nicely through the various ways in which our personal digital content is more at risk than we might think (media degradation, device and format obsolescence, and the sheer anonymity of large quantities of digital stuff). But he has a prescription for dealing with some of it, part of which I reproduce here (I hope fair use covers this, since you'll have to go to the original to read the rest!):
"However, all is not lost. There are some strategies for storing your digital archives. But you'll have to do a lot of work. You will need to start thinking like a librarian and become an active curator of your files. That means relentlessly organizing, labeling and tagging, backing up and deleting.

The first and most important thing to do is to begin deleting files. Whittle things down to the essentials. What do you really want to maintain and pass along? You must be ruthless and vigilant.

Next, develop a system for organizing files online and offline. If you're going to store stuff on removable media, like DVDs, place them in cases that have extensive labels, and index them. And don't store files like text documents or photos on propriety formats that are not widely adopted. Experts recommend photos in JPG forms and documents in PDF formats or basic text formats.

Label every file and tag them with as much information as you can. Being obsessive now will pay off in the long run. This is a lot of work, which is why you want to cull your archives as much as possible.

Once that's done, make multiple copies. You can also explore "cloud" backup services..."
Thinking like a librarian? Being an active curator of your files? Sounds like a good place to start. Interesting that he sees deleting as being an important part of remembering! We probably need better tools for the average person for a lot of this (eg tagging files in a filestore), but I suspect there's enough around for any reasonably competent researcher to use. However, laziness, forgetfulness and sheer pressure of work are our enemies here. Will we forget to do the things needed to remember?

Wednesday, 5 August 2009

Curated databases and data curation

I've been working on an article with a colleague, and came across something that's interesting, and that you might be able to help us with. There does appear to be a distinction between the way curation is used in the bio-sciences, and elsewhere. In particular, the term "curated database" tends to mean a manually constructed database that links literature to data, curated by experts who provide authority (eg see the Wikipedia definition of Biocurator). The earliest mention of the term "curated database" I can find is in the abstract (and only in the abstract) of Larsen et al (1993).

However, the terms "digital curation" and "data curation" tend to mean something different. We in the DCC say "Digital curation is maintaining and adding value to a trusted body of digital information for current and future use; specifically, we mean the active management and appraisal of data over the life-cycle of scholarly and scientific materials". This has a lot more elements of good management about it, and less of the idea of specific curators making judgements based on the literature. It has also been rather conflated with digital preservation.

The earliest reference to digital curation I can find is a report of an invitational meeting held in October 2001, oddly titled "The Digital Curation: digital archives, libraries and e-science seminar" (Beagrie and Pothen, 2001). In the meeting there was some discussion about data curation. The earliest more formal reference to data curation I can find is a technical report from Gray, Szalay et al (2002).

So my challenge is this: are there earlier references to digital (or data) curation, of the second kind?

Beagrie, N., & Pothen, P. (2001). The Digital Curation: digital archives, libraries and e-science seminar. Ariadne. http://www.ariadne.ac.uk/issue30/digital-curation/.

Gray, J., Szalay, A. S., Thakar, A. R., Stoughton, C., & vandenBerg, J. (2002). Online Scientific Data Curation, Publication, and Archiving. Redmond: Microsoft Research. http://arxiv.org/abs/cs.DL/0208012.

Larsen, N., Olsen, G. J., Maidak, B. L., McCaughey, M. J., Overbeek, R., Macke, T. J., et al. (1993). The ribosomal database project. Nucl. Acids Res., 21(13), 3021-3023. http://nar.oxfordjournals.org/cgi/content/abstract/21/13/3021

Wednesday, 15 July 2009

Digital Curation Conference deadline extended

5th International Digital Curation Conference (IDCC09)
Moving to Multi-Scale Science: Managing Complexity and Diversity.
2 – 4 December 2009, Millennium Gloucester Hotel, London, UK.
**************************************************************************
We are pleased to announce that the Paper Submission date for IDCC09 has been extended by 2 weeks to Friday 7 August 2009: http://www.dcc.ac.uk/events/dcc-2009/call-for-papers/
Remember that submissions should be in the form of a full or short paper, or a one page abstract for a poster, workshop or demonstration.
Presenting at the conference offers you the chance to:-

  • Share good practice, skills and knowledge transfer
  • Influence and inform future digital curation policy & practice
  • Test out curation resources and toolkits
  • Explore collaborative possibilities and partnerships
  • Engage educators and trainers with regard to developing digital curation skills for the future

Speakers at the conference will include:-

  • Timo Hannay – Publishing Director, Nature.com
  • Professor Douglas Kell – Chief Executive of the Biotechnology & Biological Sciences Research Council (BBSRC)
  • Dr. Ed Seidel, Director of the National Science Foundation’s Office of Cyberinfrastructure

All papers accepted for the conference will be published in the International Journal of Digital Curation

Sent on behalf of the Programme Committee –

co-chaired by Chris Rusbridge, Director of the Digital Curation Centre, Liz Lyon, Director of UKOLN and Clifford Lynch, Executive Director of the Coalition for Networked Information.

Wednesday, 10 June 2009

How can Social Bookmarking tools support community resource building?

In the DCC we are trying to work out ways that we can present tools to the community that help you to help us, to help you. The most primitive example of this would be the use of email lists: having identified some issue, we ask a question on a list, and use feedback from list members to develop our response to the issue. But we want to go further; just not sure how.

In this post I want to explore two use cases where social bookmarking tools might be helpful, and to seek advice on how to take these ideas forward. The two cases are:
  • getting input from the community on curation tools and resources worth investigating
  • extending a proposed bibliography on data curation.
In the first case, we’re looking for suggestions for quality curation resources. These could be tools of various kinds, guides, policies, templates, even standards (although we have separate ideas on the latter). We currently have a form on our web site for suggestions, but it’s long-winded and clunky, and we don’t get much input. What if we could use something like Delicious. It’s very easy to bookmark a resource with Delicious, as I’m sure you know. A couple of clicks and a few keystrokes, and you’ve bookmarked and tagged something. But how can we arrange for Delicious bookmarking to feed in to a set of resources for us to review? I wondered if asking people to use a tag such as <DCC-suggest> might work?

In the second case, I have been building an extensive bibliography of books, articles and reports relevant to research data curation, management and preservation. We can load such a bibliography onto our web site in a variety of formats, including simple web pages for reading, and downloadable BibTex, RIS or other formats. But that leaves the bibliography as a static resource, and the responsibility for maintenance and enhancement lies entirely with us. And if someone identifies a good candidate, there’s no easy way to feed it into the bibliography.

Now there are a few social bookmarking sites that are specifically oriented towards managing references, including Connotea. But I can’t work out how to use them in this way. I had a go at using Connotea a year or so ago, but have largely given up because it wasn’t very good at extracting the metadata for the kinds of resources I was bookmarking (so I had to do all the work anyway), and while I could do a download once from Connotea into the reference management tool I was then beginning to use on my desktop (a commercial product I won’t name), I couldn’t work out how to do incremental downloads. I had another poke around today, and while there clearly is some way of sharing, it didn’t feel like the simple act that social networking requires. And I couldn’t see much value in other people’s tags.

Today after reading an interesting article (Hull, Pettifer, & Kell, 2008), I experimented with Mendeley, which looked interesting. I’m not sure it works a lot better for me, for various reasons (although the metadata extraction works a bit better), but it was hard to be convinced it would be useful for this use case, given relatively low usage. I also remembered that I played with CiteULike a while ago; again I couldn’t quite work out how to use it as I want to, either personally, or in this use case.

I’m hoping that there is some way, with one or other of these tools, to load up the bibliography, maybe tagged in some way such as <data-curation>. That might allow others to find and access these resources, download the bookmarks etc. People could also presumably upload further bookmarks and tag them with the same tag, so that adds to the resources available to others. I’m not sure what can be done in this circumstance to quality-validate these resources, so that the whole bibliography remains of appropriate quality. Any ideas?


Hull, D., Pettifer, S. R., & Kell, D. B. (2008). Defrosting the Digital Library: Bibliographic Tools for the Next Generation Web. PLoS Comput Biol, 4(10), e1000204. http://dx.doi.org/10.1371%2Fjournal.pcbi.1000204

Tuesday, 2 June 2009

New JISC Research Data Management Programme

I have not been posting for over a month now, due to pressure of work (bid writing, mainly), and its curiously hard to get back into it. There's plenty of things to write about; too many, really; too hard to pick the right one. But now something has arrived that absolutely demands to be blogged: the JISC Research Data Programme, Data Management Infrastructure Call for Projects (JISC 07/09) has been released (closing data 6 August 2009, Community Briefing 6 July 2009)!

We are very excited about this. The Programme aims for 6-8 projects in English or Welsh (not Scottish or Irish, grrrr) institutions that will identify requirements to manage data created by researchers, and then will deploy a pilot data management infrastructure to address these requirements. Coupled with some projects already funded under the JISC 12/08 call (CLARION at Cambridge, Bril at Kings, EIDCSR at Oxford, Lifespan RADAR at Royal Holloway, and the Materials Data Centre at Southampton), and some work not eligible for this call, such as Edinburgh DataShare, these projects will really begin to build experience in managing research data at institutional level (or groups of institutions) in the UK.

The DCC has a key role in the Programme. Apart from mentions of the Curation Lifecycle and the Data Audit Framework, paragraph 30 of the Call says
"30. Role of the Digital Curation Centre (DCC): Bidders are invited to consult with the DCC in preparing their bids. The DCC will provide general support for this strand of activities and for the programme more broadly. This will be done by contributions to programme events as well as the current channels of information, and through its principal role as a broker for expertise and advice in the management and curation of data. Projects are encouraged to engage directly with the DCC and its programme of information exchange - for example, by contributing to the Research Data Management Forum (RDMF)."
We are preparing to deliver on this role both in the short term (during the bid-writing phase), and later during the Programme execution. I have been very pleased to work with Simon Hodson, the JISC Programme Manager for this Call. I'm sure Simon does not yet realise how much influence he will be exerting over the future of research data curation in institutions in the UK!

Wednesday, 15 April 2009

5th International Digital Curation Conference: Call for Papers

The Call for Papers for the 5th International Digital Curation Conference has just been published. With the title "Moving to Multi-Scale Science: Managing Complexity and Diversity", the conference will be held in London from 2-4 December, 2009. I believe this is THE conference for papers on advances in digital and data curation! The text of the call follows:
We invite submission of full papers, posters, workshops and demos and welcome contributions and participation from individuals, organisations and institutions across all disciplines and domains that are engaged in the creation, use and management of digital data, especially those involved in the challenge of curating data for e-science and e-research.

Proposals will be considered for short (up to 6 pages) or long (up to 12 pages) papers and also for demonstrations, workshops and posters. The full text of papers will be peer-reviewed; abstracts for all posters, workshops and demos will be reviewed by the co-chairs. Final copy of accepted contributions will be made available to conference delegates, and papers will be published in our International Journal of Digital Curation [external]. Accordingly, we recommend that you download our template and read the advice on its use.

Papers should be original and innovative, probably analytical in approach, and should present or reference significant evidence (whether experimental, observational or textual) to support their conclusions.

Subject matter could be policy, strategic, operational, experimental, infrastructural, tool-based, and so on, in nature, but the key elements are originality and evidence. Layout and structure should be appropriate for the disciplinary area. Papers should not have been published in their current or a very similar form before, other than as a pre-print in a repository.

We seek papers that respond to the main themes of the conference: multi-scale, multi-discipline, multi-skill and multi-sector, and that relate to the creation, curation, management and re-use of research data. Research data should be interpreted broadly to include the digital subjects of all types of research and scholarship (including Arts and Humanities, and all the Sciences). Papers may cover:
  • Curation practice and data management at the extremes of scale (e.g. interactions between small science and big science, or extremes of object size, numbers of objects, rates of deposit and use)
  • Challenging content: (e.g. addressing issues of data complexity, diversity and granularity)
  • Curation and e-research, including contextual, provenance, authenticity and other metadata for curation (e.g. automated systems for acquiring such metadata)
  • Research data infrastructures, including data repositories and services
  • Disciplinary and inter-disciplinary curation challenges and data management approaches, standards and norms
  • Promoting, enabling, demonstrating and characterizing the re-use of data
  • Semantically rich documents (e.g. the “well-supported article”)
  • The human infrastructure for curation (e.g. skills, careers, training and organisational support structures, careers, skills, training and curriculum)
  • Curation across academia, government, commerce and industry
  • Legal and policy issues; Creative Commons, special licences, the public domain and other approaches for re-use, and questions of privacy, consent, and embargo
  • Sustainability and economics: understanding business and financial models; balancing costs, benefits and value of digital curation
Important Dates
  • Submission of papers for peer-review: 24 July 2009
  • Submission of abstracts posters/demos/workshops: 24 July 2009
  • Notification of authors of papers: 18 September 2009
  • Notification of authors of posters/demos/workshops: 2 October 2009
  • Final papers deadline: 13 November 2009
  • Final posters deadline: 13 November 2009

Wednesday, 18 February 2009

Notes from NERC Data Management workshop 1

David Bloomer, NERC CIO (and Finance Director) opened the workshop, and talked about data acquisition, data curation, data access, and data exploitation, in the context of developing NERC Information Strategy. Apparently NERC does not currently have an Information Strategy, as the last effort was thrown out in Council. Clearly from his point of view, the issue was about working out whether data management is being done well enough, and how it can be done better within the resources available.

There were some interesting comments about licensing and charging: principles that I summarise as:
  • encouraging re-use and re-purposing,
  • all free for teaching and research,
  • that the licence and its cost depends on the USE and not the USER, and
  • that NERC should support all kinds of data users equally.
Not sure yet the full implications of this; it clearly doesn’t mean that all data is freely accessible to everyone! However, it sounds like a major improvement over recent practices, with some NERC bodies charging high prices for some of their data.

In the second session, after my talk which I have already blogged (although not the 20 minutes of panic when first the PowerPoint transferred from my Mac would not load up on Windows, and second my Mac would not recognise the external video for the projectors!), there were two short talks by Dominic Lowe from BADC on Climate Science Modelling Language, and Tim Duffy from BGS on GeoSciML.

My notes on the former are skeletal in the extreme (ie nonsense!), other than that it is an application of GML, and is based on a data model expressed with UML. However, I picked up a bit more about the latter.

The scope of GeoSciML is the scientific information linked to the geography, except the cartography. It is mainly interpretive or descriptive, and so far includes 25 internationally-agreed vocabularies. Taken 5 years or so to develop to this point. Based around GML, using UML for modelling, using OGC web services to render maps, format, query etc. Provides interoperability and perhaps a “unified” view, but does not require change to local (end) systems, nor to local schemas. Part of the claimed value of the process was exposing that they did not understand their data as well as they thought!

Wednesday, 21 January 2009

DCC Evaluation survey

We, the DCC, would like your help in evaluating our performance, and to refine our portfolio of products and services and better meet your needs. If you'd like to help then please take a moment to fill in our public survey at: <http://www.dcc.ac.uk/adding/public_survey/>

As a small token of our appreciation, we are offering one lucky entrant the chance to win an iPod nano (competition rules apply).

Tuesday, 20 January 2009

Load testing repositories

One of the issues that has worried me about moving from repositories of e-prints to repositories of data is the increased challenges of scale. Scale could be vastly different for data repositories in several dimensions, including
  • rate of deposit
  • numbers of objects
  • size of objects
  • rate of access
  • rate of change to existing objects...
Now Stuart Lewis is reporting on the first stage of the JISC-funded ROAD project, where they have load-tested a DSpace implementation (on a fairly chunky configuration), loading 300,000 digital objects of 9 MB each.

Stuart reports
  • "As expected, the more items that were in the repository, the longer an average deposit took to complete.
  • On average deposits into an empty repository took about one and a half seconds
  • On average deposits into a repository with three hundred thousand items took about seven seconds
  • If this linear looking relationship between number of deposits and speed of deposit were to continue at the same rate, an average deposit into a repository containing one million items would take about 19 to 20 seconds.
  • Extrapolate this to work out throughput per day, and that is about 10MB deposited every 20 seconds, 30MB per minute, or 43GB of data per day.
  • The ROAD project proposal suggested we wanted to deposit about 2Gb of data per day, which is therefore easily possible.
  • If we extrapolate this further, then DSpace could theoretically hold 4 to 5 million items, and still accept 2B of data per day deposited via SWORD."
They plan to repeat these tests with ePrints and FEDORA platforms.

Like all such, it's an artificial test, but it does give encouragement that DSpace could scale to handle a data repository for some tasks. I don't know if other issues would be show-stoppers or not, for something like a lab repository, but most of the scale issues seem OK.

Thursday, 8 January 2009

Digital Curation Google Group

Interesting Google Group on Digital Curation set up a month or so ago, 25 November 2008 to be exact. Brief is:
"Intended to be a collaborative space for people involved in the work of digital curation and repository development to share ideas, practices, technology, software, standards, jokes, etc."
There's been mostly techie-level discussion on various topics, including the bagit data packaging spec, whether it should include forward error correction (to cope with very long transit time, FedEx style transfers, I think), and (on a quite different level) the concept of "movage", reported by Ed Summers based on conversations at a LoC barcamp with Ryan McKinley, whoever he is, and others. The essence is:
"The only way to archive digital information is to keep it moving."
Which ties in with some thoughts of mine: the best way to preserve information is to keep using it. I don't know how to link to individual posts, but if you go look you'll find it easily enough.

One worth watching!

Saturday, 13 December 2008

What makes up data curation?

Following some discussion at the Digital Curation Conference in Edinburgh, how about this:

Data Curation comprises
  • Data management
  • Adding value to data
  • Data sharing for re-use
  • Data preservation for later re-use
Is that a good breakdown for you? Should data creation be in there? I tend to think that data creation belongs to the researcher; once created, the data immediately falls into the data management and adding value categories.

Adding value will have many elements. I’m not sure if collecting and associating contextual metadata is part of adding value, or simply good data management!

Heard at the conference: context is the new metadata (thanks Graeme Pow!).

Tuesday, 9 December 2008

Martin Lewis on University Libraries and data curation

Martin Lewis opened the second day of the International Digital Curation Conference with a provocative and amusing keynote on the possible roles of libraries in curating data. It was very early, with his presentation [large PPT] starting at 8:40 am, and the audience after the conference dinner in the splendid environs of Edinburgh Castle was unsurprisingly thin. However, absentees missed an entertaining and thought-provoking start to the day. Martin is great at the provocative remark, usually tongue in cheek; perhaps not all our visitors caught the ironic tone, a peril of speaking thus to an international audience (eg “this slide shows a spectrum of the disciplines, from the arts and humanities on this side, to the subjects with intellectual rigour over here”!).

He mentioned some common keywords when people think of libraries: conservation, preservation, permanent, but perhaps seen as intimidating, risk-averse, conservative. What can libraries do in relation to data? He promised us 8 things, but then came up with a 9th:
  1. Raise awareness of data issues,
  2. lead policy on data management,
  3. provide advice to researchers about their data management early in life cycle,
  4. work with IT colleagues to develop appropriate local data management capacity,
  5. collectively influence UK policy on research data management,
  6. teach data literacy to research students,
  7. develop KB of existing staff,
  8. work with LIS educators to identify required workforce skills (see Swan report), and
  9. enrich UG learning experience through access to research data.
Note, this list does not (currently) include much direct, practical support for data curation or preservation: no value-added data repositories, etc. Although a few libraries are venturing gently into this area, without extra support he believes that libraries are too stretched to be able to take on such a challenge. The library sector is stretched very thin, bearing in mind the major changes to accommodate the change to digital publishing, and large investments in improving support for the teaching and learning side of things, including his own libraries new building (which is leading to increasing library footfall, and increasing loans of non-digital materials, which must still be managed).

(It struck me later, by the way that this presentation was in a way quintessentially British, or at least non-American, in the extent of its attention to these non-research elements; American chief librarian colleagues have seemed to be focused on the research library and less interested in the under-graduate library.)

So bearing the resource problems in mind, Lewis took us through the Library/IT response to the English Funding Council’s Shared Services programme: the UKRDS feasibility study. He noted the scale of the data challenge, still poorly served in policy & infrastructure, the major gaps in current data services provision. University Library and IT Services in several institutions are coming under pressure to provide data-oriented services, mainly just now for storage (necessary but not sufficient for curation). He reminded us that the UK eScience core programme was a world leader; we had many reports to refer to, especially including Liz Lyon’s “Dealing with Data” report, but we are getting to the point where we need services not projects.

The feasibility study surprisingly shows low levels of use of national facilities; most data curation and sharing happens within the institution. The feasibility study identified 3 options:
  • do nothing,
  • massive central service, or
  • coordinated national service.
Despite an amusing side excursion exploring the imagined SWAT teams of a massive central service, swarming in from their attack helicopters to gather up neglected data, the last alternative is a clear winner for the study.

There was strong evidence base of gaps in data curation service provision in UK. Cost savings were hard to calculate in a new area, but were compared with the potential alternative of ubiquitous provision in universities for their own data. UKRDS would seek to embrace rather than replace existing services, but might provide them with additional opportunities. Next steps were to publish their report, which would recommend (and he hopes extract the funds to enable) a Pathfinder phase (operational, not pilot).

During questions, the inevitable from a representative of a discipline area with already well-established support (and I paraphrase): “how can you do the real business of curation, namely adding value to collected datasets, when you are at best generalists without real contact with the domain scientists necessary for curation?” To which, my response is “the only possible way: in partnership”!

Thursday, 4 December 2008

International Digital Curation Conference Keynote

I'm not sure how much I should be blogging about this conference, given that the DCC ran it, and I chaired quite a few sessions etc. But since I've got the scrappy notes, I may try to turn some of them into blog posts. I've spotted blog postings from Kevin Ashley on da blog, and Cameron Neylon on Science in the Open, so far.

Excellent keynote from David Porteous, articulate and passionate supporter of the major “Generation Scotland” volunteer-family-based study of Scottish Health (notoriously variable, linked to money, diet, and smoking, and in some areas notoriously poor). Some interesting demographic images based on Scottish demography 1911, 1951, 2001 and projected to 2031, show the population getting older as we know. It’s not only the increasing tax burden on a decreasing proportion of workers, but also the rise of chronic disease. The fantastically-named Grim Reaper’s Road Map of Mortality shows the very unequal distribution, with particular black spots in Glasgow. About half of these effects are “nurture” and about half are “nature”, but really it’s about the interplay between them.

He spoke next about genetics, sequencing and screening. Genome sequencing took 13 years and cost $3B, now a few weeks and $500K, next year down to a few $K? Moving to a system where we hope to identify individuals at risk, start health surveillance, understand the genetic effects and target rational drug development, we hope reducing bad reactions to drugs.

Because of population stability, the health and aging characteristics, and a few legal and practical issues (such as cradle-to-grave health records), it turns out that Scotland is particularly suited to this kind of study. Generation Scotland is a volunteer, family-based study (illustration from the Broons!). There are Centres in Edinburgh, Glasgow, Dundee, Aberdeen; no-one should be more than an hour’s travel away etc. Major emphasis on data collection, curation and integration through an integrated laboratory management system, in turn linking to the health service and its records. Major emphasis on security and privacy. Consent is open consent (rather than informed consent), but all have a right to withdraw from the study, and must be able to withdraw any future use of their data (only 2 out of nearly 14,000 have withdrawn so far!).

This wasn’t a talk about the details of curation, but it was an inspiring example of why we care about our data, and how, when the benefits are great enough and the planning and careful preparation are good enough, even major legal obstacles can be overcome.

Thursday, 20 November 2008

Curation services based on the DCC Curation Lifecycle Model

I’ve had a go at exploring some curation services that might be appropriate at different stages of a research project. I thought it might also be worth trying to explore curation services suggested by the DCC Curation Lifecycle Model, which I also mentioned in a blog post a few weeks ago.



The model claims that it “can be used to plan activities within an organisation or consortium to ensure that all necessary stages are undertaken, each in the correct sequence… enables granular functionality to be mapped against it; to define roles and responsibilities, and build a framework of standards and technologies to implement”. So it does look like an appropriate place to start.

At the core of the model are the data of interest. These should absolutely be determined by the research need, but there are encodings and content standards that improve the re-usability and longevity of your data (an example might be PDF/A, which strips out some of the riskier parts of PDF and which should make the data more usable for longer).

Curation service: collected information on standards for content, encoding etc in support of better data management and curation.
Curation service: advice, consultancy and discussion forums on appropriate formats, encodings, standards etc.

The next few rings of the model include actions appropriate across the full lifecycle. “Description and Representation Information” is maybe a bit of a mouthful, but it does cover an important step, and one that can easily be missed. I do worry that participants in a project in some sense “know too much”. Everyone in my project knows that the sprogometer raw output data is always kept in a directory named after the date it was captured, on the computer named Coral-sea. It’s so obvious we hardly need to record it anywhere. And we all know what format sprogometer data are in, although we may have forgotten that the manufacturer upgraded the firmware last July and all the data from before was encoded slightly differently. But then 2 RAs leave and the PI is laid up for a while, and someone does something different. You get the picture! The action here is to ensure you have good documentation, and to explicitly collect and manage all the necessary contextual and other information (“metadata” in library-speak, “representation information” when thinking of the information needed to understand your data long term).

Curation service: sharable information on formats and encodings (including registries and repositories), plus the standards information above.
Curation service: Guidance on good practice in managing contextual information, metadata, representation information, etc.

Preservation (and curation) planning are necessary throughout the lifecycle. Increasingly you will be required by your research funder to provide a data management plan.

Curation service: Data Management Plan templates
Curation service: Data Management Plan online wizards?
Curation service: consultancy to help you write your Data Management Plan?

Community Watch and Participation is potentially a key function for an organisation standing slightly outside the research activity itself. The NSB Long-lived data report suggested a “community proxy” role, which would certainly be appropriate for domain curation services. There are related activities that can apply in a generic service.

Curation service: Community proxy role. Participate in and support standards and tools development.
Curation service: Technology watch role. Keep an eye on new developments and critically obsolescence that might affect preservation and curation.

Then the Lifecycle model moves outwards to “sequential actions”, in the lifecycle of specific data objects or datasets, I guess. As noted before, curation starts before creation, at the stage called here “conceptualise: Conceive and plan the creation of data, including capture method and storage options”. Issues here perhaps covered by earlier services?

We then move on to “doing it”, called here “Create and Receive”. Slightly different issues depending on whether you are in a project creating the data, or a repository or data service (or project) getting data from elsewhere. In the former case, we’re back with good practice on contextual information and metadata (see service suggested above); in the latter case, there’s all of these, plus conformance to appropriate collecting policies.

Curation service: Guidance on collection policies?

Appraisal is one of those areas that is less well-understood outside the archival world (not least by me!). You’ve maybe heard the stories that archivists (in the traditional, physical world) throw out up to 97% of what they receive (and there’s no evidence they were wrong, ho ho). But we tend to ignore the fact that this is based on objective, documented and well-established criteria. Now we wouldn’t necessarily expect the same rejection fraction to apply to digital data; it might for the emails and ordinary document files for a project, but might not for the datasets. Nevertheless, on reasonable, objective criteria your lovingly-created dataset may not be judged appropriate for archiving, or selected for re-use. My guess is that this would often be for reasons that might have been avoided if different actions had been taken earlier on. Normally in the science world, such decisions are taken on the basis of some sort of peer review.

Curation service: Guidance on approaches to appraisal and selection.

At this point, the Lifecycle model begins to take on a view as from a Repository/Data Centre. So whereas in the Project Life Course view, I suggested a service as “Guidance on data deposit”, here we have Ingest: the transfer of data in.

Curation service: documented guidance, policies or legal requirements.

Any repository or data centre, or any long-lived project, has to undertake actions from time to time to ensure their data remain understandable and usable as technology moves forward. Cherished tools become neglected, obsolescent, obsolete and finally stop working or become unavailable. These actions are called preservation actions. We’ve all done them from time to time (think of the “Save As” menu item as a simple example). We may test afterwards that there has been no obvious corruption (and will often miss something that turns up later). But we’ve usually not done much to ensure that the data remain verifiably authentic. This can be as simple as keeping records of the changes that were made (part of the data’s provenance).

Curation service: Guidance on preservation actions
Curation service: Guidance on maintaining authenticity
Curation service: Guidance on provenance

I noted previously that storing data securely is not trivial. It certainly deserves thoughtful attention. At worst, failure here will destroy your project, and perhaps your career!

Curation service: Briefing paper on keeping your data safe?

Making your data available for re-use by others isn’t the only point of curation (re-use by yourself at a later date is perhaps more significant). You may need to take account of requirements from your funder or publisher or community on providing access. You will of course have to remain aware of legal and ethical restrictions on making data accessible. The Lifecycle Model calls these issues “Access, Use and Reuse”.

Curation service: Guidance on funder, publisher and community expectations and norms for making data available
Curation service: Guidance on legal, ethical and other limitations on making data accessible
Curation service: Suggestions on possible approaches to making data accessible, including data publishing, databases, repositories, data centres, and robust access controls or secure data services for controlled access
Curation service: provision of data publishing, data centre, repository or other services to accept, curate and make available research data.

Combining and transforming data can be key to extracting new knowledge from them. This is perhaps situation-specific, rather than curation-generic?

The Lifecycle model concludes with 3 “occasional actions”. The first of these is the downside of selection: disposal. The key point is “in accordance with documented policies, guidance or legal requirements”. In some cases disposal may be to a repository, or another project. In some cases, the data are, or perhaps must be destroyed. This can be extremely hard to do, especially if you have been managing your backup procedures without this in mind. Just imagine tracking down every last backup tape, CD-ROM, shared drive, laptop copy, home computer version etc! Sometimes, especially secure destruction techniques are required; there are standards and approaches for this. I’m not sure that using my 4lb hammer on the hard drives followed by dumping in the furnace counts! (But it sure could be satisfying.)

Curation Service: Guidance on data disposal and destruction

We also have re-appraisal: “Return data which fails validation procedures for further appraisal and reselection”. While this mostly just points back to appraisal and selection again, it does remind me that there’s nothing in this post so far about validation. I’m not sure whether it’s too specific, but…

Curation service: Guidance on validation

And finally in the Lifecycle model, we have migration. To be honest, this looks to me as if it should simply be seen as one of the preservation actions already referred to.

So that’s a second look at possible services in support of curation. I’ve also got a “service-provider’s view” that I’ll share later (although it was thought of first).

[PS what's going on with those colours? Our web editor going to kill me... again! They are fine in the JPEG on my Mac, but horrible when uploaded. Grump.]

Wednesday, 8 October 2008

DCC Curation Lifecycle Model

I have not written much in this blog about Digital Curation Centre products, but I think it’s time to remedy that, and mention some of them. In particular, I wanted to mention the DCC Curation Lifecycle model, which is attracting widespread interest. Primarily put together by Sarah Higgins with input from colleagues across the DCC and external experts, like all such models it is of course a compromise between succinctness and completeness. Sarah has run a couple of workshops on it, including one in the US at JCDL 08 in Pittsburgh, and the response was extremely positive.

I hope to mention later how we will be using it to structure information on standards, and we are expecting to use it as an entry point to the DCC web site and the DCC DIFFUSE Standards Frameworks Project. In addition, I learned only recently that the more detailed proposals for the UK Research Data Service (which went to their Steering Committee a week or so back) lean heavily on the model (not apparent from their interim report). It is also used to explain the roles of data managers and data scientists in the JISC report “Skills, Role & Career Structure of Data Scientists & Curators: Assessment of Current Practice & Future Needs” by Alma Swan and Sheridan Brown.

Here is the graphic summary of the model:



I won’t attempt here to explain it in detail, as there’s sufficient additional information on the DCC web site and the longer IJDC article about the model.

Wednesday, 24 September 2008

Data as major component of national research collaboration

This is perhaps the last of my posts resulting from conversations and presentations at the UK e-Science All Hands meeting in Edinburgh. This one relates to Andrew Treloar’s presentation on the Australian National Data Service (ANDS), and its over-arching programme, Platforms for Collaboration, part of the National Collaborative Research Infrastructure Strategy.

There are strong historical reasons why collaboration over a distance is very important for Universities and researchers in Australia. Now the Government seems to have really got the message, with this major investment programme. Andrew described how a basis of improving the basic infrastructure (network and access management) supports 3 programmes, high performance computing, collaboration services, and the Data Commons. The first two of these are equally important, but in from my vantage, I’m particularly interested in the Data Commons.

The latter is provided (or perhaps supported; it’s highly distributed) by ANDS. The first business plan for ANDS is available at http://ands.org.au/andsinterimbusinessplan-final.pdf and will run until July 2009, with ANDS itself expected to run to 2011. The vision for ANDS is in the document “Towards the Australian Data Commons”.

The vision document :
"...identifies a number of longer term objectives for data management:
  • A. A national data management environment exists in which Australia’s research data reside in a cohesive network of research repositories within an Australian ‘data commons’.
  • B. Australian researchers and research data managers are ‘best of breed’ in creating, managing, and sharing research data under well formed and maintained data management policies.
  • C. Significantly more Australian research data is routinely deposited into stable, accessible and sustainable data management and preservation environments.
  • D. Significantly more people have relevant expertise in data management across research communities and research managing institutions.
  • E. Researchers can find and access any relevant data in the Australian ‘data commons’.
  • F. Australian researchers are able to discover, exchange, reuse and combine data from other researchers and other domains within their own research in new ways.
  • G. Australia is able to share data easily and seamlessly to support international and nationally distributed multidisciplinary research teams. (p. 6) "
Andrew writes:
“ANDS has been structured as four inter-related and co-ordinated service delivery programs:
  • Developing Frameworks [Monash]
  • Providing Utilities [ANU]
  • Seeding the Commons [Monash]
  • Building Capabilities” [ANU]
Andrew also mentioned the Science and Research Strategic Roadmap Review, published just this August, which seems to centre on NCRIS. This includes the notion of a” national data fabric”, based on institutional nodes.

ARCHER, mentioned earlier, provides candidate technology for an institution participating in ANDS. As Andrew pointed out in a comment responding to my confusion in the last post: “the CCLRC [metadata] schema is the internal schema being use by ARCHER to manage all of the metadata associated with the experimental data. ISO2146 is the schema being used by ANDS to develop its discovery service”. No doubt there are many ways institutional nodes can and will be stitched together, and it will be interesting to see how this develops.

Australians sometimes bemoan the lack of an Australian equivalent of JISC. However, in this case they appear to have put together something with significant coherence in multiple dimensions. On the face of it, this is more significant than any European or US programme I have seen so far. A lot depends on the execution, but with good luck and a following wind (and a fairly strong dose of “suspension of disbelief” from researchers), this could well turn out to be a world-beating data infrastructure programme.

For comparison, I’ll try to take a look at the emerging UK Research Data Service proposals, shortly. Perhaps that has the opportunity to be even better?

Friday, 19 September 2008

A national data mandate? Australian Code for the Responsible Conduct of Research

Andrew Treloar pointed to this Code in a presentation at the e-Science All Hands meeting. All Australian Universities have signed up to the Code, which turns out to have a whole chapter on the management of research data and primary materials. It deals with the responsibilities of both institutions and researchers. Mostly it’s in the style “Each institution must have a policy on…”, but it then gets quite prescriptive on what those policies must cover. Here are some quotes, institutional responsibilities first:
“Each institution must have a policy on the retention of materials and research data.”

“In general, the minimum recommended period for retention of research data is 5 years from the date of publication.”

“Institutions must provide facilities for the safe and secure storage of research data and for maintaining records of where research data are stored.”

“Wherever possible and appropriate, research data should be held in the researcher’s department or other appropriate institutional repository, although researchers should be permitted to hold copies of the research data for their own use.”
On researcher responsibilities:
“Researchers should retain research data and primary materials for sufficient time to allow reference to them by other researchers and interested parties. For published research data, this may be for as long as interest and discussion persist following publication.”

“Research data should be made available for use by other researchers unless this is prevented by ethical, privacy or confidentiality matters.”

“Retain research data, including electronic data, in a durable, indexed and retrievable form.”
To discharge these obligations requires training in good curation practice, significant care, and appropriate infrastructure. Maybe these are regarded as yet more "un-funded mandates", to be treated on a risk assessment basis (will we be found out?). Maybe this does represent "the dead hand of compliance", as a senior colleague once phrased it. But if taken in the spirit as written, it represents a significant mandate for data curation!

I don’t know of an equivalent Code elsewhere that is so specific. The nearest US equivalent may be the Introduction to the Responsible Conduct of Research, from the Office of Research Integrity, Dept of Health & Human Services. It tends to be more a rather bland assembly of good advice, than anything prescriptive.

In the UK, the Research Integrity Office says it is “developing a code”, individual Research Councils have Ethics and related policies, while Research Councils UK is consulting in this area; their consultation closes on 24 October, 2008. It does have a section on Management and preservation of data and primary materials:
“… ensure that relevant primary data and research evidence are preserved and accessible to others for reasonable periods after the completion of the research. This is a shared responsibility between researcher and the research organisation, but individual researchers should always ensure that primary material is available to be checked. … Data should normally be preserved and accessible for not less than 10 years for any projects, and for projects of clinical or major social, environmental or heritage importance, the data should be retained for up to 20 years, and preferably permanently within a national collection, or as required by the funder’s data policy.”
The “normal” period is longer, but otherwise, still not as specific and therefore strong as the Australian Code!

Monday, 15 September 2008

David de Roure on "the new e-Science"

I was at the eScience All Hands meeting last week, and unfortunately missed a presentation by David de Roure on the New e-Science, an update on a talk he gave 10 months ago. The slides are available on Slideshare, but David has agreed I can share his summary:
"1. Increasing scale and diversity of participation
Decreasing cost of entry into digital research means more people, data, tools and methods. Anyone can participate: researchers in labs, archaeologists in digs or schoolchildren designing antimalarial drugs. Citizen science! Improved capabilities of digital research (e.g. increasing automation, ease of collaboration) incentivises this participation. "You're letting the oiks in!" people cry, but peer review benefits from scale of participation too. "Long Tail Science"

2. Increasing scale and diversity of data
Deluge due to new experimental methods (microarrays, combinatorial chemistry, sensor networks, earth observation, ...) and also (1). Increasing scale, diversity and complexity of digital material, processed separately and in combination. New digital artefacts like workflows, provenance, ontologies and lab books. Context and provenance essential for re-use, quality and trust. Digital Curation challenge!

3. Sharing
Anyone can play and they can play together. Anyone can be a publisher as well as a consumer - everyone's a first class citizen. Science has always been a social process, but now we're using new social tools for it. Evidenced by use of wikis, blogs, instant messaging. The lifecycle goes faster, we accelerate research and reduce time-to-experiment.

4. Collective Intelligence
Increasing participation means network effects through community intelligence: tagging, reviewing, discussion. Recommendation based on usage. This is in fact the only significant breakthrough in distributed systems in the last 30 years. Community curation: combat workflow decay!

5. Open Research
Publicly available data but also the open services and software tools of open science. Increasing adoption of Science Commons, open access journals, open data and linked data (formerly known as Semantic Web), PLoS, ... Open notebook science

6. Sharing Methods
Scripts, workflows, experimental plans, statistical models, ... Makes research repeatable, reproducible and reusable. Propagates expertise. Builds reputation. See Usefulchem, myExperiment.

7. Empowering researchers
Increasing facility with new tools puts the researchers in control - of their software/data apparatus and their experiments. Empowerment enables creativity and creation of new, sharable methods. Tools that take away autonomy will be resisted. Beware accidental disempowerment! Ultimately automation frees the researcher to do what they're best at, but can also be disempowering.

8. Better not perfect
Researchers will choose tools that are better than what they had before but not necessarily perfect. This force encourages bottom-up innovation in the practice of research. It opposes the adoption of over-engineered computer science solutions to problems researchers don't know they have and perhaps never will.

9. Pervasive deployment
Increasingly rich intersection between the physical and digital worlds through devices and instruments. Web-based interfaces not software downloads. Shift towards devices and the cloud. REST architecture coupling components that transcend their application.

10. Standing on the shoulders of giants
e-Science is now enabling researchers to do some completely new stuff! As the pieces become easy to use, researchers can bring them together in new ways and ask new questions. Boundaries are shifting, practice is changing. Ease of assembly and automation is essential."
The presentation is well worth looking at as well for the extra material David includes. I thought the more open and inclusive approach to e-Science (or Cyber-infrastructure) was well worth including here. The word "heroic" appears on his slides in relation to the Grid, which sums up my concerns, I think!