Showing posts with label DCC. Show all posts
Showing posts with label DCC. Show all posts

Monday, 1 March 2010

DCC: A new phase, a new perspective, a new Director

As the DCC begins its third phase today, I am delighted to announce the appointment of our new Director, Kevin Ashley, who will succeed me upon my retirement in April 2010.

Kevin Ashley has been Head of Digital Archives at the University of London Computer Centre (ULCC) since 1997, during which time his multi-disciplinary group has provided services related to the preservation and reusability of digital resources on behalf of other organisations, as well as conducting research, development and training. The group has operated the National Digital Archive of Datasets for The National Archives of the UK for over twelve years, delivering customised digital repository services to a range of organisations. As a member of the JISC's Infrastructure and Resources Committee, the Advisory Council for ERPANET, plus several advisory boards for data and archives projects and services, Kevin has contributed widely to the research information community. As a firm and trusted proponent of the DCC we look forward to his energetic leadership in this new phase of our evolution.

So far so press release. But I'd go further. I can't tell you how pleased I am with this appointment. As some readers will know, I have personally lobbied all and any potential candidates for this post since before I officially announced I was leaving. I understand we had some excellent candidates (I wasn't directly involved), more than one of whom might have made an excellent Director. But I'm particularly pleased at Kevin's appointment for several reasons: he is well engaged in the community including good connections with JISC, our major funder), he's tough enough to keep this tricky collaboration thing going, he has an excellent technical understanding, and he has great experience of actually managing this stuff in all its crusty awfulness. I particularly remember his discussion (on a visit to the Edinburgh Informatics Database Group) about issues like how best to deal with an archived dataset where they came across the characters "five" in a field defined as numeric! You can make it work or make it a record but not both...

So congratulations Kevin, and good luck!


Wednesday, 13 January 2010

Director of Digital Curation Centre: still time to apply

I’m particularly keen that there be a good slate of candidates for this post, for which applications close on Friday 15 January, 2010. The details can be found at http://www.jobs.ed.ac.uk/vacancies/index.cfm?fuseaction=vacancies.detail&vacancy_ref=3012085 (sorry about the dodgy URL; I hope it works)

The further details say

"The mission of the Digital Curation Centre is to help build capacity, capability and skills for data curation across the UK higher education research community, while supporting and promoting emerging data curation practice. It also has a key role in supporting JISC, especially its new research data management programme. Overall, the DCC is an agent for change, committed to the diffusion of best practice in the curation of digital research data across the Higher Education sector, and providing an authoritative source of advocacy, resources and guidance to the UK research community. This mission is informed by five priorities:

to identify, gather, record and disseminate curation best practice, providing access to resources, tools, training and information that will equip data practitioners to make informed decisions regarding the management of their data assets;

to facilitate knowledge exchange between those currently and newly engaged in the generation and management of digital research data;

to build and support a community of informed practitioners that has the capacity to sustain itself, with the capability to manage and curate its data appropriately;

to identify crucial and important innovations in data curation, and seek additional resources to provide them;

to support JISC, especially in its repository, preservation and data management programmes.

To achieve this, the Director must be a persuasive advocate for better curation and management of research data, on a national or international scale. Able to listen and engage with researchers and with research management, publishers and research funders, the Director will build a strong, shared vision of the changes needed, and the ability (working with others in the DCC and beyond) to mobilise the community towards that end."

So if this fits you, or you know someone that it fits, please persuade that person to apply!

By the way, although the DCC may not escape further budget cuts, like all public services in the UK, I have been told that we are funded from core JISC funding rather than capital funding, and as such Phase 3 will not be curtailed as the proposed JISC Managing Research Data programme has been.

[NOTE: This vacancy has now CLOSED; no further applications will be accepted!]

Tuesday, 5 January 2010

Digital Curation Centre User Survey 2009: Highlights

My colleague Angus Whyte has provided the following brief summary of two surveys carried out in Phases 1 and 2 of the Digital Curation Centre, in 2006 and 2009 respectively, as part of our evaluations. In retrospect, we might have done better revising the questions for the second survey rather more than we did; nevertheless I thought it worth while sharing this with you.

Angus writes:

In 2009 DCC users were surveyed, repeating a similar survey carried out in 2006. In the highlights below we draw conclusions both from the more recent results and also changes over the 3 year period. Both surveys were publicised on the DCC website and via several mailing lists, principally the DCC-Associates and (in 2009) the JISC sponsored Research-Dataman list.

Our conclusions take into account that the online questionnaire was self-completed by a self-selected group of respondents (75 in 2009 and 125 in 2006). DCC Associates (640 approx.) provided the bulk of the responses[1]. The results indicated broad patterns, relatively wide differences and consistent responses over the two surveys, even though these are not taken to be statistically representative.

Highlights

In both surveys around 90% of respondents are familiar with the term ‘digital curation’ and regard it as a critical issue within their project or unit. The DCC is consistently given as the main source of information on curation issues by around 70% of respondents, with “on the job challenges/ research” second at around 60%.

Between the two surveys there is a large jump (from 13% to 32%) in the number of respondents indicating that DCC has been “very effective” in raising awareness about digital curation, and those believing it to be “slightly effective” has correspondingly fallen from 53% to 31%.

Of a list of DCC resources, five are identified as “most helpful” by at least 1 in 5 of the 2009 survey respondents, these being (in descending order) the DCC website, Briefing Papers (of various sorts), the DCC Curation Lifecycle Model, Case Studies, and the Digital Curation Manual.

Respondents universally associate digital curation with “ensuring the long-term accessibility and re-usability of digital information”, and large majorities (around 90%) also relate it to “performing archiving activities on digital information such as selection, appraisal and retention” and “ensuring the authenticity, integrity and provenance of digital information are maintained over time”. Rather lower but still significant numbers (around 60%) associate digital curation with “managing digital information from its point of creation” and “managing risks to digital information” – although many more highlight the latter in 2009 (up to 84% from 61%).

Curation or preservation addresses risks to the respondents’ organisations with “loss of organisational memory” consistently topping their list (identified by around 75% of respondents) and “business risks” second, identified by just under half, again across both surveys.

More than two thirds indicate that their main reasons for curating and preserving digital information are its educational/research or historical value; in both years a minority cites other reasons. Similarly, the main obstacles are indicated as financial or staff resources, with around half also indicating lack of awareness or appropriate policies.

For around 40% of respondents, management and preservation of digital information has an indefinite timescale. For a further 15% or so it is “beyond the life of the project/organisation”, and similar numbers indicate these are tasks “for the life of the project/organisation”.

The 2009 survey respondents are no strangers to the ‘data deluge’, most dealing with at least 100Gb and some (7%) more than 100Tb. Overall 79% expect this to increase in the next two years, surprisingly 3% do not, while 7% do not know. Most need to manage a mixture of open and proprietary formats, and report a wide variety of formats in use, predominantly common office applications, PDF documents and multimedia formats. Curation and preservation challenges are most frequently identified with obsolete proprietary formats. Image, video, and geospatial data are also often identified as challenges, as are web sites combining these.

Respondents were also asked in 2009 about re-use, and around a third indicate that research data is re-used internally, with similar numbers offering data generated by their project/unit for re-use by others, or re-using external data.

Access issues facing research projects/units are identified in both surveys and along similar lines; intellectual property rights (e.g. copyright) is the most frequently cited issue, followed by “privacy or ethical issues”, however “embargo on research findings” is least prevalent, identified by only a fifth of respondents.

Asked about funding for curation and preservation, responses show no clear picture. Around half of 2009 respondents indicate funding is “accounted for in project or institutional budget”. A large minority have no explicit funding for curation and preservation, and where resources are available these are pooled from other funded areas (e.g. IT budget for project or organisation) or research grants. Spending on curation/preservation is less than £50,000 (for around half of those respondents who were aware of this). Around half are unsure whether spending will increase or decrease, with the remainder being evenly split.

Detailed questions and response data are available on request.

Angus Whyte, Digital Curation Centre

[1] The DCC Associates membership list includes UK data organisations, leading data curators, overseas and supranational standards agencies, and industrial/business communities. Currently research data creators are under-represented (information from registration details).


Wednesday, 9 December 2009

Director, The Digital Curation Centre (DCC)

So, time to come fully out into the open, after various coy hints over the past week or so. I'm planning to retire around the start of DCC Phase 3. Adverts are starting to appear with the following text:
"We wish to appoint a new Director to take the DCC forward into an exciting third phase, from March 2010. You must be a persuasive advocate for better management of research data on a national and international scale. Able to listen to and engage with researchers and with research management, publishers and research funders, you will build a strong, shared vision of the changes needed, working with and through the community. You should have a sound knowledge of all aspects of digital curation and preservation, an understanding of higher education structures and processes, and appropriate management skills to be the guiding force in the DCC’s progress as an effective and enduring organisation with an international reputation.

"This post is fixed term for three years. [I would suggest: in the first instance...]
Closing date: Friday 15th January 2010."
This is a great job, and I think an important one. I have been bending the ears of many people in the last couple of weeks to ask them to think of appropriate people to point this advert at (yes, I know that's rotten English; it's been a long week).

Further details will be on the University of Edinburgh's jobs web site, http://www.jobs.ed.ac.uk/ (they weren't there when I checked a few minutes ago, maybe tomorrow).

Tuesday, 8 December 2009

Leadership opportunities

Those interested in leadership in Digital and Data Curation should keep an eye on the relevant UK press and lists over the next week or so for anything of interest...

Friday, 14 August 2009

DCC web site and Linked Data

We at the DCC are in the early stages of refreshing our web site (www.dcc.ac.uk). Nothing you can see yet, but we're talking to a few consultants about what and how we can do better. The ones we have spoken to so far seem pretty clued up on content management systems, and even on web 2.0 approaches. But questions about the role of the Semantic Web or Linked Data get blank looks.

Now our web site is not and will probably never be a major source of data as facts; rather it should contain resources: often documents, sometimes tools, sometimes sharing opportunities. There definitely are facts of various kinds there (which may not sufficiently explicit yet), such as staff contact details, document metadata, event locations and times, etc. But these are a comparatively small part of the content.

Does this (or anything else) justify investment in building a web site that is based on Linked Data/Semantic Web? What advantages could we get in doing so? What advantages could our users get if we did so?

I would really like to get some views on this!

Wednesday, 10 June 2009

How can Social Bookmarking tools support community resource building?

In the DCC we are trying to work out ways that we can present tools to the community that help you to help us, to help you. The most primitive example of this would be the use of email lists: having identified some issue, we ask a question on a list, and use feedback from list members to develop our response to the issue. But we want to go further; just not sure how.

In this post I want to explore two use cases where social bookmarking tools might be helpful, and to seek advice on how to take these ideas forward. The two cases are:
  • getting input from the community on curation tools and resources worth investigating
  • extending a proposed bibliography on data curation.
In the first case, we’re looking for suggestions for quality curation resources. These could be tools of various kinds, guides, policies, templates, even standards (although we have separate ideas on the latter). We currently have a form on our web site for suggestions, but it’s long-winded and clunky, and we don’t get much input. What if we could use something like Delicious. It’s very easy to bookmark a resource with Delicious, as I’m sure you know. A couple of clicks and a few keystrokes, and you’ve bookmarked and tagged something. But how can we arrange for Delicious bookmarking to feed in to a set of resources for us to review? I wondered if asking people to use a tag such as <DCC-suggest> might work?

In the second case, I have been building an extensive bibliography of books, articles and reports relevant to research data curation, management and preservation. We can load such a bibliography onto our web site in a variety of formats, including simple web pages for reading, and downloadable BibTex, RIS or other formats. But that leaves the bibliography as a static resource, and the responsibility for maintenance and enhancement lies entirely with us. And if someone identifies a good candidate, there’s no easy way to feed it into the bibliography.

Now there are a few social bookmarking sites that are specifically oriented towards managing references, including Connotea. But I can’t work out how to use them in this way. I had a go at using Connotea a year or so ago, but have largely given up because it wasn’t very good at extracting the metadata for the kinds of resources I was bookmarking (so I had to do all the work anyway), and while I could do a download once from Connotea into the reference management tool I was then beginning to use on my desktop (a commercial product I won’t name), I couldn’t work out how to do incremental downloads. I had another poke around today, and while there clearly is some way of sharing, it didn’t feel like the simple act that social networking requires. And I couldn’t see much value in other people’s tags.

Today after reading an interesting article (Hull, Pettifer, & Kell, 2008), I experimented with Mendeley, which looked interesting. I’m not sure it works a lot better for me, for various reasons (although the metadata extraction works a bit better), but it was hard to be convinced it would be useful for this use case, given relatively low usage. I also remembered that I played with CiteULike a while ago; again I couldn’t quite work out how to use it as I want to, either personally, or in this use case.

I’m hoping that there is some way, with one or other of these tools, to load up the bibliography, maybe tagged in some way such as <data-curation>. That might allow others to find and access these resources, download the bookmarks etc. People could also presumably upload further bookmarks and tag them with the same tag, so that adds to the resources available to others. I’m not sure what can be done in this circumstance to quality-validate these resources, so that the whole bibliography remains of appropriate quality. Any ideas?


Hull, D., Pettifer, S. R., & Kell, D. B. (2008). Defrosting the Digital Library: Bibliographic Tools for the Next Generation Web. PLoS Comput Biol, 4(10), e1000204. http://dx.doi.org/10.1371%2Fjournal.pcbi.1000204

Wednesday, 21 January 2009

DCC Evaluation survey

We, the DCC, would like your help in evaluating our performance, and to refine our portfolio of products and services and better meet your needs. If you'd like to help then please take a moment to fill in our public survey at: <http://www.dcc.ac.uk/adding/public_survey/>

As a small token of our appreciation, we are offering one lucky entrant the chance to win an iPod nano (competition rules apply).

Wednesday, 8 October 2008

DCC Curation Lifecycle Model

I have not written much in this blog about Digital Curation Centre products, but I think it’s time to remedy that, and mention some of them. In particular, I wanted to mention the DCC Curation Lifecycle model, which is attracting widespread interest. Primarily put together by Sarah Higgins with input from colleagues across the DCC and external experts, like all such models it is of course a compromise between succinctness and completeness. Sarah has run a couple of workshops on it, including one in the US at JCDL 08 in Pittsburgh, and the response was extremely positive.

I hope to mention later how we will be using it to structure information on standards, and we are expecting to use it as an entry point to the DCC web site and the DCC DIFFUSE Standards Frameworks Project. In addition, I learned only recently that the more detailed proposals for the UK Research Data Service (which went to their Steering Committee a week or so back) lean heavily on the model (not apparent from their interim report). It is also used to explain the roles of data managers and data scientists in the JISC report “Skills, Role & Career Structure of Data Scientists & Curators: Assessment of Current Practice & Future Needs” by Alma Swan and Sheridan Brown.

Here is the graphic summary of the model:



I won’t attempt here to explain it in detail, as there’s sufficient additional information on the DCC web site and the longer IJDC article about the model.

Monday, 8 September 2008

OAIS revision moving forward?

Just over a year ago, in late August 2007 I wondered what was happening with the required review of the Open Archival Information System standard, which was announced in June 2006, and for which comments closed in October 2006. Well, there is at last some movement. Just recently, the DCC and the Digital preservation Coalition received notice of the "proposed dispositions, with rationale, to the suggestions which your organisation sent in response to the request for recommendations for updates to the OAIS Reference Model (ISO 14721)". According to the email from John Garrett, Chair of the CCSDS Data Archiving and Ingest Working Group,
"If you have feedback on these proposed dispositions please email them as soon as possible, and by 30th November 2008 at the latest, [...]

A revised draft of the full OAIS Reference Model is expected to be available on the Web in January 2009. There will then be a period for further comment before submission to ISO for full review."
Although a fair few of the dispositions are "No changes are planned", a large number of changes are also proposed relating to the comments made by DCC/DPC. We have not yet had a chance to review them in detail, nor even to decide yet what the mechanism for this will be. But I am very much encouraged that progress is at last being made, and that more opportunities to interact with the development of this important standard will be available, even if it has not proved possible to find out the venue where the proposed changes have been discussed!

Monday, 11 August 2008

New issue of International Journal of Digital Curation

I am very pleased to announce the publication of Volume 3, Issue 1 of the International Journal of Digital Curation, at http://www.ijdc.net/

This is the largest issue so far, including 9 peer-reviewed papers and 8 articles. My thanks to all the contributors, and to Richard Waller for his excellent editorial work.

Papers (Peer-reviewed)
  • Evolving a Network of Networks: The Experience of Partnerships in the National Digital Information Infrastructure and Preservation Program. Martha Anderson
  • Toward Distributed Infrastructures for Digital Preservation: The Roles of Collaboration and Trust. Michael Day
  • Dataset Preservation for the Long Term: Results of the DareLux Project. Eugène Dürr, Kees van der Meer, Wim Luxemburg, Ronald Dekker
  • Curation of Laboratory Experimental Data as Part of the Overall Data Lifecycle. Jeremy Frey
  • Towards a Theory of Digital Preservation. Reagan Moore
  • Challenges and Issues Relating to the Use of Representation Information for the Digital Curation of Crystallography and Engineering Data. Manjula Patel, Alexander Ball
  • Defining File Format Obsolescence: A Risky Journey. David Pearson, Colin Webb
  • Data Documentation Initiative: Toward a Standard for the Social Sciences. Mary Vardigan, Pascal Heus, Wendy Thomas
  • Moving Archival Practices Upstream: An Exploration of the Life Cycle of Ecological Sensing Data in Collaborative Field Research. Jillian C. Wallis, Christine L. Borgman, Matthew S. Mayernik, Alberto Pepe
Articles
  • The DCC / Regional eScience Collaborative Workshop. Martin Donnelly
  • The DCC Curation Lifecycle Model. Sarah Higgins
  • What to Preserve?: Significant Properties of Digital Objects. Helen Hockx-Yu, Gareth Knight
  • Recycling Information: Science Through Data Mining. Michael Lesk
  • The Fit Between the UK Environmental Information Regulations and the Freedom of Information Act. Colin Pelton, Mark Thorley
  • Review: Scholarship in the Digital Age. Chris Rusbridge
  • Meeting Curation Challenges in a Neuroimaging Group. Angus Whyte, Dominic Job, Stephen Giles, Stephen Lawrie

Thursday, 10 April 2008

4th Digital Curation Conference: Call for papers

The call is now open at http://www.dcc.ac.uk/events/dcc-2008/.
"The first day of the conference will focus on three key topics:

  • Radical sharing, new ways of doing science e.g. large scale research networks, mass collaboration, dynamic publishing tools, wikis, blogs, social networks, visualisations and immersive environments
  • Sustainability of curation
  • Legal issues including privacy, confidentiality and consent, intellectual property rights and provenance
"The second day of the conference will be dedicated to research and development and will feature peer-reviewed papers in themed parallel sessions."
There are several main themes:

  • Research Data Infrastructures (covering research data across all disciplines)
  • Curation and e-Research
  • Sustainability: balancing costs and value of digital curation and preservation
  • Disciplinary and Inter-disciplinary Curation Challenges
  • Challenging types of content
  • Legal Issues
  • Capacity Building
Key dates:
  • Submission of papers for peer-review: 25 July 2008
  • Submission of abstracts poster/demos/workshops for peer-review: 25 July 2008
  • Notification of authors: 19 September 2008
  • Final papers deadline: 14 November 2008
  • Submission of poster PDFs: 14 November 2008
We need good papers, so please get your thinking caps on!

Thursday, 20 March 2008

Legacy document formats

On the O'Reill XML blog, which I always read with interest (particularly in relation to the shenanigins over OOXML and ODF standardisation), Rick Jelliffe writes An Open Letter to Microsoft, IBM/Lotus, Corel and others on Lodging Old File Formats with ISO. He points out that
"Corporations who were market leaders in the 1980s and 1990s for PC applications have a responsibility to make sure that documentation on their old formats are not lost. Especially for document formats before 1990, the benefits of the format as some kind of IP-embodying revenue generator will have lapsed now in 2008. However the responsibility for archiving remains.

"So I call on companies in this situation, in particular Microsoft, IBM/Lotus, Corel, Computer Associates, Fujitsu, Philips, as well as the current owners of past names such as Wang, and so on, to submit your legacy binary format documentation for documents (particularly home and office documents) and media, to ISO/IEC JTC1 for acceptance as Technical Specifications.[...] Handing over the documentation to ISO care can shift the responsibility for archiving and making available old documentation from individual companies, provide good public relations, and allow old projects to be tidied up and closed."
This is in principle a Good Idea. However, ISO documents are not Open Access; the specifications Rick refers to would benefit greatly from being Open. They would form vitally important parts of our effort to preserve digital documents. Instead of being deposited in ISO, they should be regarded as part of Representation Information for those file types, and deposited in a variety (more than one, for safety's sake) of services such as PRONOM at The National Archive in the UK, the proposed Harvard/Mellon Global Digital Format Registry, the Library of Congress Digital Preservation activity or the DCC's own Registry/Repository of Representation Information.

Friday, 14 December 2007

Murray-Rust on Digital Curation Conference day 2

Peter Murray-Rust has written two blog posts (here and here; I'm not sure if those are permanent URL's...) about day 2 of the International Digital Curation Conference in Washington DC. Thanks, Peter.

In the first post, he began:
"There is a definitely an air of optimism in the conference - we know the tasks are hard and very very diverse but it’s clear that many of them are understood."
He then picked up on Carole Goble's presentation on workflows. Here are a few random extracts from Peter's random jottings (his description):
"The great thing about Carole is she’s honest. Workflows are HARD. They are expensive. There are lots of them. Not of them does exactly what you want. And so on. [PMR: We did a lot of work - by our standards - on Taverna but found it wasn’t cost-effective at that stage. Currently we script things and use Java. Someday we shall return.]"
"myExperiment.org. A collaborative site for workflows. You can go there and find what you want (maybe) and find people to talk to. “- bazaar for workflows, encapsulated objects (EMO) single WFs or collections, chemistry data with blogged log book, encapsulatd experimental objects Open Linked Data linked initiative…"
"Scientists do not collaborate - scientists would rather share a toothbrush [...] than gene names (Mike Ashburner)
who gets the credit? - who is allowed to update?. Changing metadata rather than data. Versioning. Have to get credit and reputation managed. Scientitsts are driven by money, fame, reputation, fear of being left behind"
"Annotations are first class citizens"
His second post covers Jane Hunter and Kwok Cheung's presentation on compound document objects (CDOs):
"Increasing pressure to share and publish data while maintaining competitiveness.
Main problem lack of simple tools for recording, publishing, standards.
What is the incentive to curate and deposit? What granularity? concern for IP and ownership"
"Current problems with traditional systems - little semantic relationship, little provenance, little selectivity, interactivity , flexibility and often fixed rendering and interfaces. No multilevel access. either all open or all restricted
usually hardwired presentation"
"Capture scientific provenance through RDF (and can capture events in physical and digital domain)
Compound Digital Objects - variable semantics, media, etc.
Typed relationships within the CDOs. (this is critical)"
"SCOPE [the tool Jane & Kwok have developed is] a simplified tool for authoring these objects. Can create provenance graphs. Infer types as much as possible. RSS notification. Comes with a graphical provenance explorer."
Thanks again, Peter!

Thursday, 13 December 2007

Murray-Rust on Digital Curation Conference day 1

Peter Murray-Rust has blogged on our conference in Washington DC:
Overall impressions - optimistic spirit with some speakers being very adventurous about what we can and should do.
He paid particular attention to the national perspective from Australia:
"…Rhys Francis (Australia) . One of the most engaging presentations. Why support ICT? It will change the world. Systems can make decisions, electronic supply chains. humans cannot keep pace. We do not need people to process information.

Who owns the data, and the copyright = physics says, who cares? If it’s good, copy it. Else discard it. Storage is free. [PMR: let’s try this in chemistry…]

What to do about data is a harder question than how to build experiements

FOUR components to infrastructure. I’ve seen this from OZ before… data, compute, interoperation, access. Must do all 4. -

And he also highlighted the importance of domain knowledge in preference to institutional repositories (I go along with that).
And I thought this was interesting, relating to Tim Hubbard's presentation:
"particular snippet: Rfam (RNA family) is now in Wikipedia as primary resource with some enhancement by community annotation. So we are seeing Wikipedia as the first choice of deposition for areas of scholarship."
So it looks like things are going well!

Wednesday, 12 December 2007

DCC Forum: Comment on Conference best paper

There is a Conference theme over on the DCC Forum. However, despite its RSS feed, it doesn't appear to be quite in the blogosphere, so I thought I would post some quotes from it here. Bridget Robinson announced the best paper selection (it gets a star spot in the programme):
"The title of best peer-reviewed paper has been awarded to "Digital Data Practises and the Long Term Ecological Research Programme" by Helena Karasti (Department of Information Processing, University of Oulu in Finland), Karen S Baker (Scripps Institution of Oceanography, University of California, USA) and Katharina Schleidt (Department of Data and Diagnoses, Umweltbundesant, Austria)."
Angus Whyte commented:
"Can't make it to Washington, but looking forward to seeing this paper. I'm familiar with the 'Enriching the Notion of Data Curation in e-Science' paper by the same authors (plus Eija Halkola) in the journal Computer Supported Cooperative Work last year [*}. (doi:10.1007/s10606-006-9023-2); and some of Helena Karasti's previous work to integrate ethnographic and participatory design methods in CSCW. Recently colleagues and I at DCC began case studies for the SCARP project, which I hope can draw lessons from their work to document "the inter-related and continuously changing nature of technology, data care, and science conduct" as they put it.
* Volume 15, Number 4 / August, 2006

If there are any coments here I will re-post a summary back on the Forum!

Tuesday, 11 December 2007

3rd International Digital Curation Conference

The 3rd International Digital Curation Conference is getting underway about now with a pre-conference reception at the National Museum of the American Indian in Washington DC. We are very pleased that it is fully booked, and the Programme Committee have put together an excellent programme. I'm hoping that the presentations will be mounted on the web site very quickly[*], and also that some of those present will be blogging the event (if you do, folks, please tag with IDCC3 so that I can find them).

Yes, as you may have guessed, unfortunately I can't be there either this year, which right now makes me hopping mad (madder than a cut snake, as my erstwhile compatriots used to say).
So have a good one, guys!!!

[*Added later: perhaps slides could be added to Slideshare? Again with tag IDCC3...]

Monday, 10 December 2007

Two fundamentally different views on data curation

A few months ago, we interviewed two scientists for two quite different posts in the DCC. Both were from a genomics background, and I was very struck by how strongly held, but from my point of view how narrow, was their view of curation. As a result, I’m beginning to realise that there are two fundamentally different approaches to data curation

The Wikipedia definition of biocurator was quoted by Judith Blake, the keynote speaker [slides] at the 2nd International BioCuration Meeting in San Jose:
“A biocurator is a professional scientist who collects, annotates, and validates information that is disseminated by biological and model organism databases. The role of a biocurator encompasses quality control of primary biological research data intended for publication, extracting and organizing data from original scientific literature, and describing the data with standard annotation protocols and vocabularies that enable powerful queries and biological database inter-operability. Biocurators communicate with researchers to ensure the accuracy of curated information and to foster data exchanges with research laboratories.

Biocurators (also called scientific curators, data curators or annotators) have been recognized as the ‘museum catalogers of the Internet age’ [Bourne & McEntyre]”
In essence, this suggests that curation (let’s call it biocuration for now) is the construction of authoritative annotations linking significant objects (eg genes) with evidence about them in the literature and elsewhere. And this is indeed the sort of thing you see in genomics databases.

In other parts of science, I believe curation is not so much about constructing annotation on objects, but rather about caring for arbitrary datasets (adding descriptions, information about how to use them, transferring them to longer term homes, clarifying conditions of use and provenance information, but yes including annotations, etc), so that they can be used and re-used, by the originators or others, now or in the future. This approach incorporates elements of digital preservation, but is not solely defined by “long term” in the way that digital preservation tends to be. Clearly we will find cases where we need to curate datasets which themselves result from biocuration!

The problem is that both terms are well-embedded in their respective communities. I don’t think either of my two interviewees had any idea how much broader our concept was than theirs. We need to watch out for the inevitable misunderstandings that can arrive from this clash of meanings! However, both communities are likely to continue to use the terms in their own ways (although perhaps this term biocurator, relatively new to me, is an effort to disambiguate).

Wednesday, 29 August 2007

IJDC again

At the end of July I reported on the second issue of the International Journal of Digital Curation (IJDC), and asked some questions:
"We are aware, by the way, that there is a slight problem with our journal in a presentational sense. Take the article by Graham Pryor, for instance: it contains various representations of survey results presented as bar charts, etc in a PDF file (and we know what some people think about PDF and hamburgers). Unfortunately, the data underlying these charts are not accessible!

"For various reasons, the platform we are using is an early version of the OJS system from the Public Knowledge Project. It's pretty clunky and limiting, and does tend to restrict what we can do. Now that release 2 is out of the way, we will be experimenting with later versions, with an aim to including supplementary data (attached? External?) or embedded data (RDFa? Microformats?) in the future. Our aim is to practice what we may preach, but we aren't there yet."

I didn't get any responses, but over on the eFoundations blog, Andy Powell was taking us to task for only offering PDF:
"Odd though, for a journal that is only ever (as far as I know) intended to be published online, to offer the articles using PDF rather than HTML. Doing so prevents any use of lightweight 'semantic' markup within the articles, such as microformats, and tends to make re-use of the content less easy."
His blog is more widely read than this one, and he attracted 11 comments! The gist of them was that PDF plus HTML (or preferably XML) was the minimum that we should be offering. For example, Chris Leonard [update, not Tom Wilson! See end comments] wrote:
"People like to read printed-out pdfs (over 90% of accesses to the fulltext are of the pdf version) - but machines like to read marked-up text. We also make the xml versions availble for precisely this purpose."
Cornelius Puschmann [update, not Peter Sefton] wrote:
"Yeah, but if you really want semantic markup why not do it right and use XML? The problematic thing with OJS (at least to some extent) is/was that XML article versions are not the basis for the "derived" PDF and HTML, which deal almost purely with visuals. XML is true semantic markup and therefore the best way to store articles in the long term (who knows what formats we'll have 20 years from now?). HTML can clearly never fill that role - it's not its job either. From what I've heard OJS will implement XML (and through it neat things such as OpenOffice editing of articles while they're in the workflow) via Lemon8 in the future."
Bruce D'Arcus [update, not Jeff] says:
"As an academic, I prefer the XHTML + PDF option myself. There are times I just want to quickly view an article in a browser without the hassle of PDF. There are other times I want to print it and read it "on the train."

"With new developments like microformats and RDFa, I'd really like to see a time soon where I can even copy-and-paste content from HTML articles into my manuscripts and have the citation metadata travel with it."
Jeff [update, not Cornelius Puschmann] wrote:
"I was just checking through some OJS-based journals and noticed that several of them are only in PDF. Hmmm, but a few are in HTML and PDF. It has been a couple of years since I've examined OJS but it seems that OJS provides the tools to generate both HTML and PDF, no? Ironically, I was going to do a quick check of the OJS documentation but found that it's mostly only in PDF!

"I suspect if a journal decides not to provide HTML then it has some perceived limitations with HTML. Often, for scholarly journals, that revolves around the lack of pagination. I noticed one OJS-based journal using paragraph numbering but some editors just don't like that and insist on page numbers for citations. Hence, I would be that's why they chose PDF only."
I think in this case we used only PDF because that was all our (old) version of the OJS platform allowed. I certainly wanted HTML as well. As I said before, we're looking into that, and hope to move to a newer version of the platform soon. I'm not sure it has been an issue, but I believe HTML can be tricky for some kinds of articles (Maths used to be a real difficulty, but maybe they've fixed that now).

I think my preference is for XHTML plus PDF, with the authoritative source article in XML. I guess the workflow should be author-source -> XML -> XHTML plus PDF, where author-source is most likely to be MS Word or LaTeX... Perhaps in the NLM DTD (that seems to be the one people are converging towards, and it's the one adopted by a couple of long term archiving platforms)?

But I'm STILL looking for more concrete ideas on how we should co-present data with our articles!

[Update: Peter Sefton pointed out to me in a comment that I had wrongly attributed a quote to him (and by extension, to everyone); the names being below rather than above the comments in Andy's article. My apologies for such a basic error, which also explains why I had such difficulty finding the blog that Peter's actual comment mentions; I was looking in someone else's blog! I have corrected the names above.

In fact Peter's blog entry is very interesting; he mentions the ICE-RS project, which aims to provide a workflow that will generate both PDF and HTML, and also bemoans how inhospitable most repository software is to HTML. He writes:
"It would help for the Open Access community and repository software publishers to help drive the adoption of HTML by making OA repositories first-class web citizens. Why isn't it easy to put HTML into Eprints, DSpace, VITAL and Fez?

"To do our bit, we're planning to integrate ICE with Eprints, DSpace and Fedora later this year building on the outcomes from the SWORD project – when that's done I'll update my papers in the USQ repository, over the Atom Publishing Protocol interface that SWORD is developing."
So thanks again Peter for bringing this basic error to my attention, apologies to you and others I originally mis-quoted, and I look forward to the results of your efforts! End Update]

OAIS review: what's happening?

In June 2006, there was an announcement:
In compliance with ISO and CCSDS procedures, a standard must be reviewed every five years and a determination made to reaffirm, modify, or withdraw the existing standard. The “Reference Model for an Open Archival Information System (OAIS)” standard was approved as CCSDS 650.0-B-1 in January 2002 and was approved as ISO standard 14721 in 2003. While the standard can be reaffirmed given its wide usage, it may also be appropriate to begin a revision process. Our view is that any revision must remain backward compatible with regard to major terminology and concepts. Further, we do not plan to expand the general level of detail. A particular interest is to reduce ambiguities and to fill in any missing or weak concepts. To this end, a comment period has been established.
Comments were required by 30 October 2006. The Digital Preservation Coalition and the Digital Curation Centre ran a joint workshop on 13 October in Edinburgh, and as a result submitted joint comments. Some of these comments were minor but some were quite significant. There are 14 general recommendations, and many detailed updates and clarifications were suggested.

Supporters of OAIS (and I am one) often make a great play about how open their process was. In that spirit, I have tried to find out what is happening, and to take part. I understood there was to be an open process, with a wiki and telecons, to decide in the first place whether the standard needs revision, and if so to revise it. Despite numerous attempts, I cannot find out what the current state is, or where or how this is taking place. Does anyone know?

This is a very important standard, and it needs revision to make it more useful in today's environment. We need to make it as good as we can.