Thursday, 20 March 2008

Migration on Request: OpenOffice as a platform?

Following on from my previous post relating to legacy formats, I was thinking again about the problems of dealing with documents in those formats. For some, the answer lies in emulation and perpetual licences of those original software packages, but for me that just doesn't cut the mustard. I won't have access to those packages, but I might want access to the documents. Some of them for example, might be the PowerPoint 4 presentations created on a predecessor to the Macintosh that I use now, but which are un-readable with my current PowerPoint software (I CAN get at them by copying them to a colleague's Windows machine; her version of PowerPoint has input filters unavailable on my Mac).

So I want some form of migration. In the example above, this is known as "Save as"!

However, I know that every time I do migration I introduce some sort of errors. So if I migrate from those PowerPoint 4 files to today's PowerPoint, and then from today's to tomorrow's PowerPoint, and then from tomorrow's to the next great thing, I will introduce cumulative errors whose impact I will only be able to assess at some horribly cringe-making moment, like in the middle of a presentation using a host's machine. So the best way to do migration is to start from the original file and migrate to today's version. Always. It's nuts for Microsoft to drop old file format support from its software (at least from this pint of view).

This approach of migrating from original version to today's version is called Migration on Request, and was described in a paper by Mellor, Wheatley and Sergeant back in 2002 (I referred to it earlier), but the idea hasn't caught on much. They had some other great ideas, like writing the migration tool in a specially portable version of C with all the nasty bits removed, called C--.

I have wondered from time to time however, for that class of documents we call Office Documents (word processing, spreadsheets, presentations), whether tacking onto an open source project which has a strong developer community might be a better approach. Something like OpenOffice. I'm not sure how many file formats this already supports (always growing, I guess, but Chapter 3 of the "Getting Started" documentation lists the following:
Microsoft Word 6.0/95/97/2000/XP) (.doc and .dot)
Microsoft Word 2003 XML (.xml)
Microsoft WinWord 5 (.doc)
StarWriter formats (.sdw, .sgl, and .vor)
AportisDoc (Palm) (.pdb)
Pocket Word (.psw)
WordPerfect Document (.wpd)
WPS 2000/Office 1.0 (.wps)
DocBook (.xml)
Ichitaro 8/9/10/11 (.jtd and .jtt)
Hangul WP 97 (.hwp)
.rtf, .txt, and .csv
... which is not a bad list (just the word processing bit, too)... and maybe extended in more up to date versions. For interest their FAQs have a question "Why does OpenOffice not support the file format my application uses?"
"There may be several reasons, for example:
  • The file formats may not be open and available.
  • There may not be enough developers available to do the work (either paid or volunteer).
  • There may not be enough interest in it.
  • There may be reasonable, available workarounds."
Making legacy file formats more open was the subject of my previous post, and I guess we have to wait and see. But there are plenty of legacy word processing formats not on that list (Samna, for example, later to evolve into Lotus Word Pro, as well as formats for obsolete computers like the Atari, such as the German word processor SIGNUM, supposedly very good for mathematical formulae). What about earlier version of MS Word? Wikipedia lists a bunch of word processors; there must be many documents in obscure locations in these formats.

With a concerted effort, we could gradually build OpenOffice input filters for these obsolete document types, thus brining them into the preservable digital world. And this is an effort that could bring in that extraordinary community of enthusiasts who do so much to build document converters and other kinds of software, so much ignored by the digital preservation community!

Legacy document formats

On the O'Reill XML blog, which I always read with interest (particularly in relation to the shenanigins over OOXML and ODF standardisation), Rick Jelliffe writes An Open Letter to Microsoft, IBM/Lotus, Corel and others on Lodging Old File Formats with ISO. He points out that
"Corporations who were market leaders in the 1980s and 1990s for PC applications have a responsibility to make sure that documentation on their old formats are not lost. Especially for document formats before 1990, the benefits of the format as some kind of IP-embodying revenue generator will have lapsed now in 2008. However the responsibility for archiving remains.

"So I call on companies in this situation, in particular Microsoft, IBM/Lotus, Corel, Computer Associates, Fujitsu, Philips, as well as the current owners of past names such as Wang, and so on, to submit your legacy binary format documentation for documents (particularly home and office documents) and media, to ISO/IEC JTC1 for acceptance as Technical Specifications.[...] Handing over the documentation to ISO care can shift the responsibility for archiving and making available old documentation from individual companies, provide good public relations, and allow old projects to be tidied up and closed."
This is in principle a Good Idea. However, ISO documents are not Open Access; the specifications Rick refers to would benefit greatly from being Open. They would form vitally important parts of our effort to preserve digital documents. Instead of being deposited in ISO, they should be regarded as part of Representation Information for those file types, and deposited in a variety (more than one, for safety's sake) of services such as PRONOM at The National Archive in the UK, the proposed Harvard/Mellon Global Digital Format Registry, the Library of Congress Digital Preservation activity or the DCC's own Registry/Repository of Representation Information.

Tuesday, 18 March 2008

Novartis/Broad Institute Diabetes data

Graham Pryor spotted an item on the CARMEN blog, pointing to a Business Week article (from 2007, we later realised) about a commercial pharma (Novartis) making research data from its Type 2 Diabetes studies available on the web. This seemed to me an interesting thing to explore (as a data person, not a genomics scientist), both for what it was, and for how they did it.

I could not find a reference to these data on the Novartis site, but I did find a reference to a similar claim dating back to 2004, made in the Boston Globe and then in some press releases from the Broad Institute in Cambridge, MA, referring to their joint work with Novartis (eg initial announcement, first results and further results). The first press release identified David Altshuler as the PI, and he was kind enough to respond to my emails and point me to their pages that link to the studies and to the results they are making available.

Why make the data available? The Boston Globe article said "Commercially, the open approach adopted by Novartis represents a calculated gamble that it will be better able to capitalize on the identification of specific genes that play a role in Type 2 diabetes. The firm already has a core expertise in diabetes. Collaborating on the research will give its scientists intimate knowledge of the results."

The Business Week article said "...the research conducted by Novartis and its university partners at MIT and Lund University in Sweden merely sets the stage for the more complex and costly drug identification and development process. According to researchers, there are far more leads than any one lab could possibly follow up alone. So by placing its data in the public domain, Novartis hopes to leverage the talents and insights of a global research community to dramatically scale and speed up its early-stage R&D activities."

Thus far, so good. Making data available for un-realised value to be exploited by others is at the heart of the digital curation concept. There are other comments on these announcements that cynically claim that the data will have already been plundered before being made accessible; certainly the PIs will have first advantage, but there is nothing wrong with that. The data availability itself is a splendid move. It would be very interesting to know if others have drawn conclusions from the data (I did not see any licence terms, conditions, or even requests such as attribution, although maybe this is assumed as scientific good practice in this area).

Business Week go on to draw wider conclusions:
"The Novartis collaboration is just one example of a deep transformation in science and invention. Just as the Enlightenment ushered in a new organizational model of knowledge creation, the same technological and demographic forces that are turning the Web into a massive collaborative work space are helping to transform the realm of science into an increasingly open and collaborative endeavor. Yes, the Web was, in fact, invented as a way for scientists to share information. But advances in storage, bandwidth, software, and computing power are pushing collaboration to the next level. Call it Science 2.0."
I have to say I'm not totally convinced about the latter phrase. Magazines like Business Week do like buzz-words like Science 2.0, but so far comparatively little science is affected by this kind of "radicalsharing". Genomics is definitely one of the poster children in this respect, but the vast majority of science continues to be lab or small group based with an orientation towards publishing results as papers, not data.

So what have they made available? There are 3 diabetes projects listed:
  1. Whole Genome Scan for Type 2 Diabetes in a Scandinavian Cohort
  2. Family-based linkage scan in three pedigrees with extreme diabetes phenotypes
  3. A Whole Genome Admixture Scan for Type 2 Diabetes in African Americans
The second of these does not appear to have data available online. The 3rd project has results data in the form of an Excel spreadsheet, with 20 columns and 1294 rows; the data appear relatively simple (a single sheet, with no obvious formulae or Excel-specific issues that I could see), and could probably have been presented just as easily as CSV or another text variant. There's a small amount of header text in row 2 that spans columns, plus some colour coding, that may have justified the use of Excel. Short to medium term access to these data should be simple.

The first project shows two different types of results, with a lot more data: Type 2 Diabetes results and Related Traits results. The Type 2 Diabetes results comprise a figure in JPEG or PDF, plus data in two forms: a HTML table of "top single-marker and multi-marker results", and a tab-delimited text file (suitable for analysis with Haploview 4.0) of "all single-marker and multi-marker results". These data are made available both as the initial release of February 2007, and an updated release from March 2007. There is a link to Instructions for using the results files, effectively short-hand instructions for feeding the data into Haploview and doing some analyses on them. The HTML table is just that; data in individual cells are numbers or strings, without any XML or other encoding. There are links to entries in NCBI, HapMap and Ensembl, however.

The Related Traits results also come in an initial release (also February 2007) and an updated release from September 2007. The results again have a summary, a table this time but still in JPEG or PDF form. The detailed results are more complex; there is a HTML table of traits in 4 groups (Glucose, Obesity, Lipid and Blood Pressure), and for each trait (eg Fasting Glucose) up to 4 columns of data. The first column is a description of the trait as a PDF, the next is a link to a HTML Table of Top Single Marker Results for Association, the next is a link to a text Table of All Single Marker Results for Association, and the last is a link to a text table of Phenotype summary statistics by genotype (both these have the same format as above, although the latter has different columns).

It seems clear that there is a lot of data here; how useful they are to other scientists is not for me to judge. Certainly a scientist looking through these pages could form judgments on the usefuleness and relevance of these data to his or her work. There's not much to help a robot looking for science data from the Internet. I'm not sure what form such information might take, although there are examples in Chemistry. Perhaps the data cells should be automatically encoded according to a relevant ontology, so that the significance of the data travels with them. Possibly microformats or RDFa could have (or come to have) some relevance. However, both the HTML and text formats are very durable, (more so than the Excel format for project 3) and should be easily accessible (or transformed into later forms) at least as long as the Broad Institute wishes to continue to make them available.

Friday, 7 March 2008

Data, repositories and Google

In a post last year, Peter Murray Rust criticised DSpace as a place to keep data:
"The search engines locate content. Try searching for NSC383501 (the entry for a molecule from the NCI) and you’ll find: DSpace at Cambridge: NSC383501

"But the actual data itself (some of which is textual metadata) is not accessible to search engines so isn’t indexed. So if you know how to look for it through the ID, fine. If you don’t you won’t. [...]

"So (unless I’m wrong and please correct me), deposition in DSpace does NOT allow Google to index the text that it would expose on normal web pages. [...]

"If this is true, then repositing at the moment may archive the data but it hides it from public view except to diligent humans. So people are simply not seeing the benefit of repositing - they don’t discover material though simple searches."
Peter isn't often wrong, but in this case it was clear from comments to his post that Google does normally index DSpace content, not just the metadata. There were a couple of reasons for the effects Peter saw, but the key one related to the nature of the data. Jim Downing wrote, for example:
"Not sure what to tell you about your ChemML files. Possibly Google doesn’t know what to do with them and doesn’t try?

"That’s my understanding - interestingly, if you lie about the MIME type, Google does index CML (here, for example)."
The data Peter refers to is Chemical Markup Language data in a file with extension .cml. My Mac does not know what it is, and I guess no more does Google… unless perhaps you tell Google that it’s text, as Jim Downing seemed to be suggesting in his comment (I’m not sure this constitutes lying, more selective use of the truth). I can open CML files in my text editor, fine, although of course to process them into something chemically interesting, I would need some additional software or plugins… Here's a chunk of that file [sorry, tried to include some XML here but Blogger swallowed it up]...

There's real data here [trust me: INCHI and SMILE at least, plus bond strengths etc] that could be indexed but isn't. The point is, surely, that this would be just as much a problem if the repository was simply a filestore full of CML files, which is how data is often made available. But unlike the filestore, there is usually some useful metadata in the repository which can assist data users (ie people, in this case); in a filestore, this is either absent, encoded in filenames, or in some conventional place such as README.TXT where it's relation to the actual data file is problematic).

So: in the first place, Google et al are unlikely to index data, particularly unusual data types. And in the second place, repositories encourage metadata, which does get indexed. So from this point of view at least, a repository may provide better exposure for your data (and hence more data re-use) than simply making the files web-accessible.

This doesn't mean that current, library-oriented repositories are yet fit for purpose for science data! Far from it...

Sunday, 27 January 2008

Repositories for the people?

I have been doing some thinking over the past couple of months about the role of repositories in digital curation, and it appears that others have as well. Dorothea Salo (Digital Repository Services Librarian at George Mason University) wrote a series of fascinating posts on her Caveat Lector blog, illustrating through fictional personae the dilemmas faced by many of the players in science, in academic and non-academic management in the fictional University of Achaea. The players are:
Among many interesting aspects that she highlights, the following quote is perhaps the most scary (partly because it is so often true):
"Dr. Troia’s basketology data, which are unique and could not be recreated if they were to disappear, live on the computer in her office. This computer is not to her knowledge backed up. Dr. Troia doesn’t want to put data on the department’s shared network drive, because she isn’t sanguine about its security, and her data are vital to her professional advantage, not to be pawed over by just anyone. Some of her older data are in a file format her current software can’t open; Dr. Troia shrugs about that—it’s just how software works, and she has a workaround (though a tedious and annoying one) for any file she absolutely must get into."
There is other evidence for this problem. For example, the Survey Report from the SToRe project notes:
"2.3.6 On a matter related to access and protection, the apparently common practice of storing unique and original research data on the hard drives of PCs and laptops, which seemed to be especially prevalent in chemistry and the biosciences, is a cause for concern."
Anecdotally, many of those hard drives were (and are) inadequately backed up. (Mea culpa here, having lost a months worth of data on a lost laptop. It's perhaps worth mentioning in passing here that convenient, reliable backup for a small department or group of independent-minded researchers each using their own technology approach such as Windows, Mac, Linux, desktop or laptop etc, can be surprisingly hard to organise. But I guess that's a topic for another day.)

In a related post, Dorothea writes:
"It’s really hard to put yourself in the mindset of “the user” when you’re trying to put a piece of software together. Working from standards or other specifications, that’s easy; the worst you’ll run into is linguistic ambiguity. But “the user”? Who the heck is “the user”? And that’s the point. There’s no such animal as “the user.” There’s various sorts of people who will be using your software. Persona development is an attempt to ditch “the user” in favor of a portrait that actually resonates, a concrete mental image that one can build a product or service design around."
Dorothea went down the path of creating personae because, although we think we know about researchers, "We have not been designing repositories for a typical, representative faculty member." She goes on to list a number of claimed repository advantages, that just don't cut any ice with Dr Troia. But:
"What Dr. Troia does need, and would immediately admit that she needs, is secure, networked, backed-up, maybe even version-controlled storage with access controls. Not (I cannot say this strongly enough) archival storage, because she needs to use it for in-progress work. But storage that had archival bolted onto the side with nice pointers about what to archive and when and how, that she might well use. Goes double if the system makes it easier for her to send her work to journals, or comply with data-retention or funder open-access requirements.

Do our repositories make any of that easier? Do they hell. They make it harder, and they don’t sweeten the pot by solving Dr. Troia’s other problems. No wonder all the Dr. Troias out there don’t use them."
This ties in with a comment from Peter Murray Rust in a much more recent post:
"I shall argue that the conventional model where information is “put” into “repositories” is the wrong design - certainly for data. Repositories have to be part of the scientific process - the key person is the scientist. "
The repository has to be in the science workflow for it to be used properly!

Saturday, 5 January 2008

More posts on the 3rd Digital Curation Conference

Somewhat belatedly, I found a series of posts blogged during the 3rd International Digital Curation Conference in Washington last month by K-State Libraries Conference Reports (more than one individual). The posts included (with some highly selective personal choice extracts):

Digital Curation Conference: a general comment
"I just spent four days hearing over and over again, however, how thousands of scientists in many countries make use of the very same technologies used by your average teenager--namely, blogs, wikis, Facebooky things, etc.--to do essential work and collaborate with their peers. The head of Web publishing at the Nature Publishing Group, certainly no lightweight or frivolous organization, proudly states that they're creating social networking tools built directly on the Facebook model, and sees in blogging part of the future of scientific communication. Scientist after scientist stood up and showed how wikis and blogs are used for essential communication and collaboration, how some labs have gone nearly entirely over to a social collaborative model."
Digital Curation Conference: Surveying Bloggers' Perspectives
Digital Curation Conference: Sustaining Engineering Informatics
"[Joshua Lubell] then outlined the three access scenarios (the three Rs): reference, reuse, and rationale. The first is simply the ability to view and visualize the engineering data. Reuse is what it sounds like (STEP [ISO 10303] supports this), namely, taking the data and modifying or reengineering it. Rationale is the ability to display information such as construction history, design intent, etc., which go beyond the design itself. He compared this to the chess column in a newspaper, where, in addition to a snapshot of the board at some point in the game, you get a description of the moves required to reach that state. Loosely put, it's the set of 'why' questions that can arise from a design. STEP does not include this rationale piece, which is the point behind his work."

Digital Curation Conference: Moving Archival Practices Upstream
Digital Curation Conference: Day Two Keynote
"Carol Goble, U of Manchester

"[n.b.- This talk was the highlight of DCC for me. Rather than highlighting, again, only the challenges without really offering solutions, she showed concrete examples and tools, in pretty good detail.]"
Digital Curation Conference: Towards Distributed Infrastructures
Digital Curation Conference: Day One Closing Plenary
"Could it be that we are well enough funded to be too comfortable with our traditional roles? Rick Luce asked this question near the end of his first day closing keynote. Otherwise, his talk was a fairly standard review of what's going on and what needs to be done to solve some of the pressing issues, but this question struck me as unique. He's right, I think. Our funding is sufficient to continue operating much as we have for a long time; sure, making incremental changes, but never really taking the great leap forward to stop doing most of what we do now and really take on some major new challenges."

Digital Curation Conference: Sustainable Access to the Records of Science
"... liked to say Fedora, too, as did many people at the CNI meeting. Clearly, it's the flavor of the week, and I sensed a lot of uncertainty among library leaders at CNI who did not yet have a Fedora installation at their library. "We need to move to Fedora now" appears to be the current mantra. That's a bit amusing. Sure, Fedora is great, and can do many wonderful things, but so can a lot of other platforms and solutions, who were either popular way back when (2005, gasp) or have yet to gain traction but are on the horizon. What I've seen and heard in the last three days convinces me that my gut feeling about Fedora is not incorrect: if you have a crew of developers, it might work well for you, but if you lack the commitment to hire and hold such a crew, Fedora is not for you."
Digital Curation Conference: National Perspectives
"In something of an aside, [Rhys Francis] pointed out that computer science began 60 years ago, communications (in the network sense) about 15-20 years ago, and he thinks that something significant and as yet unnamed and vague happened about three years ago. He said we're now living in the data deluge, and in future decades we'll look back and have a name for what is happening now. For me, that's both an exciting and somewhat unsettling notion, since those working in these areas during periods of inception and definition tend to look rather silly to their successors (don't we laugh at the notion of typing catalog cards, after all?), only much later to be recognized for their efforts and innovations.

"In his opinion, there are four facets for collaboration: data, compute, interoperation, access. One must do all four, not one or two or three, which he noted is a key message to an audience largely consisting of data managers of one sort or another.

"One of his closing questions was whether data is actually infrastructure."
David Rosenthal was also at the conference, and was also taken with Carole Goble's day 2 keynote, as he noted in his blog:
"She's a great speaker, with a vivid turn of phrase, and you have to like a talk about science on the web in which a major example is VivaLaDiva.com, a shoe shopping site."
David was interested in Carole's myexperiment implementation, and his blog is worth reading for other insights into that. But I also liked this
"The emerging world of web services is the big challenge facing digital preservation. Her talk was a wonderful illustration both of why this is an important problem, in that much of reader's experience of the Web is already mediated by services, and why the barriers to doing so are almost insurmountable."
(I think "doing so" here means preserving the reader's experience in this environment.)

10 Downing St on AHDS...

How exciting to receive an email from 10 Downing St this morning! It tells me that "The Prime Minister's Office has responded to [the AHDS] petition and you can view it here"

This refers to the AHRC's decision to cease funding the AHDS, which this blog discussed earlier. I signed the petition, the details of which were:
"On 11 May 2007, Professor Phillip Esler, Chief Executive of the AHRC, wrote to University Vice-Chancellors informing them of the Council's decision to withdraw funding from the AHDS after eleven years. The AHDS has pioneered and encouraged awareness and use among Britain's university researchers in the arts and humanities of best practice in preserving digital data created by research projects funded by public money. It has also ensured that this data remains publically available for future researchers. It is by no means evident that a suitable replacement infrastructure will be established and the AHRC appears to have taken no adequate steps to ensure the continued preservation of this data. The AHDS has also played a prominent role in raising awareness of new technologies and innovative practices among UK researchers. We believe that the withdrawal of funding for this body is a retrogade step which will undermine attempts to create in Britain a knowledge economy based on latest technologies. We ask the Prime Minister to urge the AHRC to reconsider this decision."
The Prime Minister's response is:
"Thank you for your e-petition.

"Government policy is that decisions on such services are a matter for the Research Council concerned.

"The Government expects that such decisions would only be taken after a careful review of the service in question. We are aware that the Arts and Humanities Research Council (AHRC) only took the decision to cease funding the Arts and Humanities Data Service (AHDS) after detailed consideration. The Council concluded at its March 2007 meeting that having regard both to the performance and cost of AHDS, this was not something they should continue to fund. There are no grounds for the Government seeking to ask the Council to reconsider that decision."
Well, I guess those who wrote and those who signed the petition didn't expect much more than this. But I think I have yet to see any "suitable replacement infrastructure" and the AHRC still "appears to have taken no adequate steps to ensure the continued preservation of this data".