Showing posts with label life-cycle. Show all posts
Showing posts with label life-cycle. Show all posts

Thursday, 20 November 2008

Curation services based on the DCC Curation Lifecycle Model

I’ve had a go at exploring some curation services that might be appropriate at different stages of a research project. I thought it might also be worth trying to explore curation services suggested by the DCC Curation Lifecycle Model, which I also mentioned in a blog post a few weeks ago.



The model claims that it “can be used to plan activities within an organisation or consortium to ensure that all necessary stages are undertaken, each in the correct sequence… enables granular functionality to be mapped against it; to define roles and responsibilities, and build a framework of standards and technologies to implement”. So it does look like an appropriate place to start.

At the core of the model are the data of interest. These should absolutely be determined by the research need, but there are encodings and content standards that improve the re-usability and longevity of your data (an example might be PDF/A, which strips out some of the riskier parts of PDF and which should make the data more usable for longer).

Curation service: collected information on standards for content, encoding etc in support of better data management and curation.
Curation service: advice, consultancy and discussion forums on appropriate formats, encodings, standards etc.

The next few rings of the model include actions appropriate across the full lifecycle. “Description and Representation Information” is maybe a bit of a mouthful, but it does cover an important step, and one that can easily be missed. I do worry that participants in a project in some sense “know too much”. Everyone in my project knows that the sprogometer raw output data is always kept in a directory named after the date it was captured, on the computer named Coral-sea. It’s so obvious we hardly need to record it anywhere. And we all know what format sprogometer data are in, although we may have forgotten that the manufacturer upgraded the firmware last July and all the data from before was encoded slightly differently. But then 2 RAs leave and the PI is laid up for a while, and someone does something different. You get the picture! The action here is to ensure you have good documentation, and to explicitly collect and manage all the necessary contextual and other information (“metadata” in library-speak, “representation information” when thinking of the information needed to understand your data long term).

Curation service: sharable information on formats and encodings (including registries and repositories), plus the standards information above.
Curation service: Guidance on good practice in managing contextual information, metadata, representation information, etc.

Preservation (and curation) planning are necessary throughout the lifecycle. Increasingly you will be required by your research funder to provide a data management plan.

Curation service: Data Management Plan templates
Curation service: Data Management Plan online wizards?
Curation service: consultancy to help you write your Data Management Plan?

Community Watch and Participation is potentially a key function for an organisation standing slightly outside the research activity itself. The NSB Long-lived data report suggested a “community proxy” role, which would certainly be appropriate for domain curation services. There are related activities that can apply in a generic service.

Curation service: Community proxy role. Participate in and support standards and tools development.
Curation service: Technology watch role. Keep an eye on new developments and critically obsolescence that might affect preservation and curation.

Then the Lifecycle model moves outwards to “sequential actions”, in the lifecycle of specific data objects or datasets, I guess. As noted before, curation starts before creation, at the stage called here “conceptualise: Conceive and plan the creation of data, including capture method and storage options”. Issues here perhaps covered by earlier services?

We then move on to “doing it”, called here “Create and Receive”. Slightly different issues depending on whether you are in a project creating the data, or a repository or data service (or project) getting data from elsewhere. In the former case, we’re back with good practice on contextual information and metadata (see service suggested above); in the latter case, there’s all of these, plus conformance to appropriate collecting policies.

Curation service: Guidance on collection policies?

Appraisal is one of those areas that is less well-understood outside the archival world (not least by me!). You’ve maybe heard the stories that archivists (in the traditional, physical world) throw out up to 97% of what they receive (and there’s no evidence they were wrong, ho ho). But we tend to ignore the fact that this is based on objective, documented and well-established criteria. Now we wouldn’t necessarily expect the same rejection fraction to apply to digital data; it might for the emails and ordinary document files for a project, but might not for the datasets. Nevertheless, on reasonable, objective criteria your lovingly-created dataset may not be judged appropriate for archiving, or selected for re-use. My guess is that this would often be for reasons that might have been avoided if different actions had been taken earlier on. Normally in the science world, such decisions are taken on the basis of some sort of peer review.

Curation service: Guidance on approaches to appraisal and selection.

At this point, the Lifecycle model begins to take on a view as from a Repository/Data Centre. So whereas in the Project Life Course view, I suggested a service as “Guidance on data deposit”, here we have Ingest: the transfer of data in.

Curation service: documented guidance, policies or legal requirements.

Any repository or data centre, or any long-lived project, has to undertake actions from time to time to ensure their data remain understandable and usable as technology moves forward. Cherished tools become neglected, obsolescent, obsolete and finally stop working or become unavailable. These actions are called preservation actions. We’ve all done them from time to time (think of the “Save As” menu item as a simple example). We may test afterwards that there has been no obvious corruption (and will often miss something that turns up later). But we’ve usually not done much to ensure that the data remain verifiably authentic. This can be as simple as keeping records of the changes that were made (part of the data’s provenance).

Curation service: Guidance on preservation actions
Curation service: Guidance on maintaining authenticity
Curation service: Guidance on provenance

I noted previously that storing data securely is not trivial. It certainly deserves thoughtful attention. At worst, failure here will destroy your project, and perhaps your career!

Curation service: Briefing paper on keeping your data safe?

Making your data available for re-use by others isn’t the only point of curation (re-use by yourself at a later date is perhaps more significant). You may need to take account of requirements from your funder or publisher or community on providing access. You will of course have to remain aware of legal and ethical restrictions on making data accessible. The Lifecycle Model calls these issues “Access, Use and Reuse”.

Curation service: Guidance on funder, publisher and community expectations and norms for making data available
Curation service: Guidance on legal, ethical and other limitations on making data accessible
Curation service: Suggestions on possible approaches to making data accessible, including data publishing, databases, repositories, data centres, and robust access controls or secure data services for controlled access
Curation service: provision of data publishing, data centre, repository or other services to accept, curate and make available research data.

Combining and transforming data can be key to extracting new knowledge from them. This is perhaps situation-specific, rather than curation-generic?

The Lifecycle model concludes with 3 “occasional actions”. The first of these is the downside of selection: disposal. The key point is “in accordance with documented policies, guidance or legal requirements”. In some cases disposal may be to a repository, or another project. In some cases, the data are, or perhaps must be destroyed. This can be extremely hard to do, especially if you have been managing your backup procedures without this in mind. Just imagine tracking down every last backup tape, CD-ROM, shared drive, laptop copy, home computer version etc! Sometimes, especially secure destruction techniques are required; there are standards and approaches for this. I’m not sure that using my 4lb hammer on the hard drives followed by dumping in the furnace counts! (But it sure could be satisfying.)

Curation Service: Guidance on data disposal and destruction

We also have re-appraisal: “Return data which fails validation procedures for further appraisal and reselection”. While this mostly just points back to appraisal and selection again, it does remind me that there’s nothing in this post so far about validation. I’m not sure whether it’s too specific, but…

Curation service: Guidance on validation

And finally in the Lifecycle model, we have migration. To be honest, this looks to me as if it should simply be seen as one of the preservation actions already referred to.

So that’s a second look at possible services in support of curation. I’ve also got a “service-provider’s view” that I’ll share later (although it was thought of first).

[PS what's going on with those colours? Our web editor going to kill me... again! They are fine in the JPEG on my Mac, but horrible when uploaded. Grump.]

Monday, 17 November 2008

Project data life course

This blog post is an attempt to explore the “life course” of an arbitrary small to medium research project with respect to data resources involved in the project. (I want to avoid the term life cycle, since we use this in relation to the actual data.)

This seems like a useful exercise to attempt, even if it has to be an over-generalisation, as understanding these interactions may help to define services that support projects in managing and curating their data better, and a number of such services are suggested here. This is about “Small Science”; James M. Caruthers, a professor of chemical engineering at Purdue University has claimed ‘Small Science will produce 2-3 times more data than Big Science, but is much more at risk’ (Carlson, 2006). Large and very large projects tend to have very particular approaches, and probably represent significant expertise; I’ll not presume to cover them.

Is this generalisation way off the mark? I’d really like to know how to improve it, accepting that it is a generalisation!

The overall life course of a project (I suggest) is roughly:
  • pre-proposal stage (when ideas are discussed amongst colleagues, before a general outline of a proposal is settled on), followed by
  • the proposal stage (constrained by funder’s requirements, and by the need to create a workable project balancing expected science outcomes, resources available, and papers and other reputational advantages to be secured). If the project is funded, there follows
  • an establishment phase, when the necessary resources are acquired, configured and tested, (which overlaps and merges with)
  • an execution phase when the research is actually done, and
  • a termination phase as the project draws to a close.
  • Somewhere during this process, one or more papers will be produced based on the research understandings gained, linked to the data collected and analysed.
My message here is that curation is most impacted by decisions in the first 3 phases, even though it requires attention throughout!

Pre-proposal stage
At the pre-proposal stage, you and your potential partners will mostly be concerned with working out if your research idea is likely to be feasible, how and with roughly what resources. During these conversations, you will need to identify any external or already existing data resources that will be required to enable the project to succeed. You should also start to identify the data that the project would create, or perhaps update, and how these data will support the research conclusions. In doing so, attention must be paid to resource constraints, not necessarily financial, but also legal and ethical issues affecting use of the data. However, discussions are more likely to be generic rather than focused! Thus begins the impact on curation…

Curation service: 5-minute curation introduction, with pointers to more information…
Curation service: Help Desk service. Ask a Curator?

Proposal stage
Apart from the obvious focus on the detailed plan for the research itself, this stage is particularly constrained by funder requirements. Increasingly, these requirements include a data management plan. This is unlikely to be a problem for large projects, which will already have spent some time (often lots of time) considering data management issues, but may well be a new requirement for smaller projects. However, forcing attention to be paid to data management is definitely a Good Thing! I’ll have a closer look at data management plans in a later blog post.

Curation service: Data Management Plan templates
Curation service: Data Management Plan online wizards?
Curation service: consultancy to help you write your Data Management Plan?

Key issues that the project should address, and which should be reflected in the data management plan, include the standards and encodings for analysing, managing and storing data. While these must be determined by a combination of local requirements, such as the availability of licences, skills and experience, and the local IT environment, they must take account of community norms and standards. While in some cases maverick approaches can lead to breakthrough science, they can also lead to silo data, inaccessible and not re-usable (sometimes not re-usable even by their creators, after a short while). This stage begins to have a significant effect on curation, as the decisions that are fore-shadowed here will have major effects later on. Repeat after me: curation begins before creation!

Curation service: Guidance on managing data for re-use in the medium to long term
Curation service: registries and repositories to support re-use, sharing and preservation
Curation service: pointers to relevant standards, vocabularies etc.

[Note, we have concerns about the scalability of the latter, which is why the DCC DIFFUSE service is being re-modelled to be more participative.]

Establishment stage
It’s during the establishment phase that the real work is in sight. In the case of a well-established laboratory or research centre, this may merely be “yet another project” to be fitted into well-known procedures, perhaps with some adaptations. But from my observation and anecdotally, there is often a high degree of individual idiosyncratic decision-making taking place here; the focus is often on “getting down to the research”. The consequences of this for curation may be very significant, but are unlikely to be apparent until much later.

The key curation aim is to set the data management plan into operation. This requires ensuring access to external data required by the project (an activity that should have been started, at least in principle, ahead of project approval). While you’re thinking of access, rights and ownership, negotiating and documenting data ownership and access agreements for the data created in the project would be a smart idea (too often ignored, or left until problems begin to arise). Of course, it’s not just rights that are at issue, but also responsibilities, particularly where data have privacy and ethical implications.

Curation service: Briefing papers on legal, ethical and licence issues.

Meanwhile a critical step is establishing the IT infrastructure. While the focus is likely to be on acquiring and testing the data acquisition and analysis equipment, software, capabilities and workflows, in practice the apparently background activities of ensuring adequate data storage, sensible disaster recovery procedures (or at the very least, proper data backup and recovery), and thoughtful data curation processes will be critical for later re-use.

Curation service: Briefing paper on Keeping your data safe?
Curation service: tools for various tasks, eg data audit, risk assessment

It’s worth remembering that there are different kinds of data storage. While you will often be most concerned with temporary or project-lifetime storage, you may need to think about storage with greater persistence. There may be a laboratory or institutional or subject repository, external database or data service that will be critical to your project; negotiating with the curation service providers associated with these services should be an important early step. In effect, right at the start you are planning for your exit! But more on this another time…

Curation service: Guidance on data repositories for projects?
Curation service: thought-stimulating discussion papers, blog posts and case studies, journal, conference, discussion fora…

Project execution
Through the lifetime of your project, your focus will rightly be on the research itself. But you should also take time out to test your disaster recovery procedures, and check on your quality assurance. You care about the quality of your materials and syntheses (or your algorithms, or your documentary sources), you need to care as much about the quality of your data, both as-observed and as-processed. You’ll be collecting a lot of data; if some support your hypothesis and some contradict it, you’ll need to ensure you are collecting (perhaps in laboratory notebooks or equivalent) sufficient provenance and context to justify the selection of the former rather than the latter.

Not sure if there is a curation service here, or if these are good practice research guides for your particular domain… but I guess the same could be said for standards and vocabularies. Question is whether domains HAVE good practice guides for data quality?

As the project goes forward, issues may arise in your original data management plan, and you should take the time to refine it (and document the changes), and your associated curation processes. If you did take idiosyncratic decisions in haste at project establishment, at some point you may need to review them. It’s a warning sign when a staffing change is followed by proposals for major architectural changes!

Finally, at some stage you should test any external deposit procedures built into your plan.

Curation service: Guidance on data deposit…

I don’t think I’m claiming that this is the limit of curation issues during project execution; rather that here the course of projects is so varied that generalisations are even less justifiable than elsewhere. Nevertheless, the analysis here should be strengthened, so input would be helpful!

Project termination stage
As the project comes towards an end, you enter another risk phase. Firstly, most management focus is likely to be on the next project (or the one after that!). Secondly, if employment is fixed term, tied to the particular project, then research staff will have been looking for other posts as the end of the project draws near, and the project may start to look denuded (and the data management and informatics folk may be the most marketable, leaving you exposed). And of course, the final reports etc cannot be written until the project is finished, by which time the project’s staffing resources will have disappeared. It’s the PI’s job to ensure these final project tasks are properly completed, but it’s not surprising that there are problems sometimes!

Curation service: Guidance on (or assistance with) evaluation

During this final phase, you must identify any data resources of continuing value (this should have occurred earlier, but does need re-visiting), and carry out the plans to deposit them in an appropriate location, together with the necessary contextual and provenance information. Easy, hey? Perhaps not: but failure here could have a negative impact on future grants, if research funders do their compliance checking properly.

Curation service: Guidance on appraisal, Guidance on data deposit.

Finally, given that data management plans are currently in their infancy, it’s worth commenting on the plan outcomes in your project final report!

Writing papers
This is not really a project phase as such, as often several papers will be produced at different stages of the project. The key issues here are:
  • Include supplementary data with your paper where possible
  • Ensure data embedded in the text are machine readable
  • Cite your data sources
  • Ensure supportive data are well-curated and kept available (preferably accessible; this is where a repository service may come in useful).
I’ll write more on this in a separate post.

Curation service: Proposals for data citation
Curation service: Suggestions for the Well-supported article

So: the message is that curation should be a constant theme, albeit as background during most of the project execution. But the decisions taken during pre-proposal, proposal and establishment phases will have a big effect on your ability to curate your data, which may affect your research results, and will certainly affect the quality and re-usability of the data you deposit.

Carlson, S. (2006, June 23). Lost in a Sea of Science Data: Librarians are called in to archive huge amounts of information, but cultural and financial barriers stand in the way. The Chronicle of Higher Education, 52, 35.

Wednesday, 8 October 2008

DCC Curation Lifecycle Model

I have not written much in this blog about Digital Curation Centre products, but I think it’s time to remedy that, and mention some of them. In particular, I wanted to mention the DCC Curation Lifecycle model, which is attracting widespread interest. Primarily put together by Sarah Higgins with input from colleagues across the DCC and external experts, like all such models it is of course a compromise between succinctness and completeness. Sarah has run a couple of workshops on it, including one in the US at JCDL 08 in Pittsburgh, and the response was extremely positive.

I hope to mention later how we will be using it to structure information on standards, and we are expecting to use it as an entry point to the DCC web site and the DCC DIFFUSE Standards Frameworks Project. In addition, I learned only recently that the more detailed proposals for the UK Research Data Service (which went to their Steering Committee a week or so back) lean heavily on the model (not apparent from their interim report). It is also used to explain the roles of data managers and data scientists in the JISC report “Skills, Role & Career Structure of Data Scientists & Curators: Assessment of Current Practice & Future Needs” by Alma Swan and Sheridan Brown.

Here is the graphic summary of the model:



I won’t attempt here to explain it in detail, as there’s sufficient additional information on the DCC web site and the longer IJDC article about the model.

Tuesday, 29 May 2007

Life cycle approach

Digital Curation is maintaining and adding value to a trusted body of digital information over the life cycle of scholarly and scientific materials, for current and future use. It is our belief in the DCC that the curation of digital data requires this whole of life approach. Critical decisions on the curation of data are taken before the data are even created, often at the time the associated project is conceived, or funding is sought. This is not least because curation requires resources that must be allowed for within the work plan. It is increasingly clear that for any project involving data of value, you should provide a data management plan within the project proposal (NSF, 2007).

Digital curation includes good management of data for current purposes, and also in many cases the preservation of those data for the long term. Long term preservation is not necessarily an essential part of curation in all cases, although it is usually a desirable aspect (subject to appraisal and selection decisions). So we can think of curation as having two important components, which we can label “data publication”, for the process of making current data available for use by other contemporaries, and “data preservation”, for the process of making those data available for future users.

Data publication recognises that more and more “reference” works are migrating into the digital domain as curated databases, and that increasingly these are data (or sometimes combinations of data and text) rather than pure text. There are interesting questions on when and how a dataset is “published” such as those raised by Bryan Lawrence in the linked post, but I’ll skip those for now! Such reference datasets can change quite frequently, including the correction or deletion of information as well as the addition of new. The requirements for integrity and stability versus the need for change to promote accuracy bring special problems. During development of the resource, if you are interested in the long term, you will have to ensure that contextual and other information needed for preservation are gathered. This may have to happen even before you make firm decisions about preservation!

As your dataset stabilises, and in particular as it comes out of current use, it may be eligible for long term preservation. This is not an automatic choice, as resources are currently spread too thinly to preserve everything. It will be important in any data management plan to identify candidate archives for preservation that serve the appropriate “designated community” (perhaps your scientific discipline).

The locus of preservation remains an issue. Where applicable and stable, a discipline-oriented repository should ensure that selected data are curated for the long term by domain experts (unfortunately the recent disturbing decision by the AHRC to end funding of the AHDS provides a worrying precedent). The choice of repository will be important in deciding in the first case what associated information is required, and later in treating the data at critical stages in the life cycle, such as changes in the technological infrastructure or the designated community.

Should there be no discipline-oriented repository, an institutional or other local repository may be appropriate… but may be quite un-prepared to deal with data (on a recent search using OpenDOAR, only 5 institutional repositories in the UK claimed to include data, and some of those claims were false)! Keep at them…

In both cases, ensuring that users in your designated community can easily understand the meaning or information content of the data is critical. The Open Archival Information System (CCSDS, 2002) has useful recommendations for supporting the long term understandability of information, and is applicable to digital curation for that reason, although at the moment we are short of good examples of some of its concepts, and it does require some interpretation!


CCSDS (2002) Reference Model for an Open Archival Information System (OAIS). IN CCSDS (Ed.), NASA.
NSF (2007) Cyberinfrastructure Vision for 21st Century Discovery. Arlington, Virginia, National Science Foundation.