Showing posts with label Preservation formats. Show all posts
Showing posts with label Preservation formats. Show all posts

Tuesday, 28 July 2009

Rosenthal at Sun-PASIG in Malta

I was very pleased to hear David Rosenthal reprise his CNI keynote on digital preservation for the Sun-PASIG meeting in Malta, a few weeks ago now. David is a very original thinker and careful speaker. I’ve fallen into the trap before of mis-remembering him, and then arguing from my faulty version. I even noted two tweets made contemporaneously with his talk, that misquoted him and changed the meaning subtly (see below). Luckily, David has made his CNI presentation available in an annotated version on his blog, so I hope I don’t make the same mistake again.

If you were not able to hear this talk, please go read that blog post. David has some important things to say, pretty much all of which I agree strongly with. No real surprise there, as part of the talk at least echoes concerns I expressed in the “Excuse Me…” Ariadne article (Rusbridge, 2006), which on reflection was probably influenced by earlier meetings with David among others.

So here’s the highly condensed version: Jeff Rothenberg was wrong in his famous 1995 Scientific American article (Rothenberg, 1995). The important digital preservation problems for society are not media degradation or media obsolescence or format obsolescence, because important stuff is online (and more or less independent of media), and widely used formats no longer go obsolescent the way they used to when Jeff wrote the article. The important issue is money, as collecting all we need will be ruinously expensive. Every dollar we spend on non-problems (like protecting against format obsolescence) doesn’t go towards real problems.

And if you are so imbued with conventional preservation wisdom as to think that summary is nonsense, but you haven’t read the blog post, go read it before making up your mind!

David concludes:

"Practical Next Steps

Everyone - just go collect the bits: Not hard or costly to do a good enough job, Please use Creative Commons licenses

Preserve Open Source repositories: Easy & vital: no legal, technical or scale barriers

Support Open Source renderers & emulators
Support research into preservation tech: How to preserve bits adequately & affordably? How to preserve this decade's dynamic web of services? Not just last decade's static web of pages"
So what are the limitations of this analysis? My quick summary from a research data viewpoint:

Lots of important/valuable stuff is not online

Quite a lot of this stuff is not readable with common, open-source-compatible software packages

We need to keep contextual metadata as well as the bits for a lot of this stuff… and yes, we do need to learn how to do this in a scalable way.

David clearly concentrates on the online world:

Now, if it is worth keeping, it is on-line

Off-line backups are temporary”

However, it’s worth remembering Raymond Clarke’s point in my earlier post from PASIG Malta about the cost advantages of offline. Particularly in the research data world, there is a substantial set of content that exists off-line, or perhaps near-line. Some of the Rothenberg risks still apply to such content. Let’s leave aside for the moment that parallels to the scenario that Rothenberg envisages continue to exist: scholars’ works encoded in obsolete digital media are starting to be ingested in archives. But more pressingly, some research projects report that their university IT departments discourage them from using enterprise backup systems for research data, for reasons of capacity limitations. So these data often exist in a ragbag collection of scarcely documented offline media (or may even be not backed up at all). In Big Science, data may be better protected, being sometimes held in large hierarchical storage management systems. A concern I have heard from the managers of such large systems is that the time needed to migrate their substantial data holdings from one generation of storage to the next can approximate the life of the system, ie several years. And clearly such systems are more exposed to risk.

Secondly, David’s comments about format obsolescence apply specifically to common formats. He says “Gratuitous incompatibility is now self-defeating”, and “Open Source renderers [exist] for all major formats” with “Open Source isn't backwards incompatible”. But unfortunately there are examples where there are valuable resources that remain at risk. There are areas with valuable content not accessible with Open Source renderers (eg engineering and architectural design). There are many cases in research where critical analysis codes are written by non-experts, with poor version control, poorly documented. And even in the mainstream world, format obsolescence can still occur in minority formats, for all sorts of reasons, including bankruptcy, but also including sheer bad design of early versions.

Finally, I’m sure David didn’t really mean “just keep the bits”. Particularly in research, but in many other areas as well, important contextual data and metadata are needed to understand the preserved data, and to demonstrate its authenticity. The task of capturing and preserving these can be the hardest part of curating and preserving the data, precisely because those directly involved need less of the context.

Oh, that double mis-quote? Talking of the difficulty of engaging with costly lawyers, David said “1 hour of 1 lawyer ~ 5TB of disk [-] 10 hours of 1 lawyer could store the academic literature”. One tweet reported this as “Lawyer effects; cost of 10 lawyer hours could save entire academic literature!” and the other as “10 hours of a lawyer's time could preserve the entire academic literature”. See what I mean? Neither save nor preserve mean the same as store!

Overall, David does a great job, in his presentation, blog post and other writings, in reminding us not to blindly accept but to challenge preservation orthodoxy. Put simply, we have to think for ourselves.

Rothenberg, J. (1995). Ensuring the longevity of digital documents. Scientific American, 272(1), 42. http://search.ebscohost.com/login.aspx?direct=true&db=buh&AN=9501173513&site=ehost-live

(yes, that URL IS the "permanent URL according to Ebsco!)

Rusbridge, C. (2006). Excuse Me... Some Digital Preservation Fallacies? Ariadne from http://www.ariadne.ac.uk/issue46/rusbridge/.

Tuesday, 1 July 2008

Responses to RAW versus TIFF: image-related

This post summarises image-related responses to the “RAW versus TIFF” post made originally by Dave Thompson of the Wellcome Library. The key elements of Dave’s post were whether we should be archiving images using RAW (which contained more camera and exposure-related data) or TIFF (which was more standardized, likely to be accessible for longer, and had more available utilities). A subsidiary question on whether we should archive both is greatly affected by cost; responses related to this element are summarized in a separate post. As I mentioned before, these responses came on the semi-closed DCC-Associates email list. [I forgot to mention that Maureen Pennock re-posted Dave's question to this blog with some remarks of her own, and we also received a useful response which you can view there.]

One comment [on the DCC-Associates list] made by many was that JPEG2000 should be considered. Alan Morris of Morris and Ward asked if it was an ISO standard, to which Shuichi Iwata (former President of CODATA and Professor at the University of Tokyo) responded:
“Yes, it is. Organized by Joint Photographic Experts Group of ISO. It is used widely for compression of images at different levels.

http://www.jpeg.org/
James Reid of EDINA, University of Edinburgh, mentioned that “the DPC report on JPEG 2000 has relevance”. He also expressed concern:
“RAW images tend to have proprietary and sensor specific aspects that makes them less than ideal for a generic preservation format - they are of course lossless but there would be an implicit overhead in deriving sensor specific options to make preservation possible (and these may not be publicly available due to proprietary interests)...”
Larry Murray of the Public Record Office of Northern Ireland mentioned how essential RAW is to him as a keen amateur photographer, allowing him to manipulate images. He added:
“James is quite correct in his comment but personally I think the decision to save the RAW images is one of available resources. The RAW images give the data owner the opportunity to control so many of the characteristics of the original image without data loss or integrity and the ability to bulk process said changes means serious consideration must be given to their value.

For other reference on Open RAW formats see http://www.openraw.org/
Chris Puttick, CIO of Oxford Archaeology asked:
“RAW is the original - would you store a scan of a document as the primary method of preserving the document? I think RAW + [a common image format] is the best option. It does of course require greater storage space, but storage is cheap and getting cheaper every day.

Software such as DigiKam has libraries that read pretty much all the RAW formats including rarer ones such as that produced by the Foveon sensor (Sigma, Polaroid). Any additional documentation needed for RAW should be easily prised out of the sensor manufacturers if the right pressure is brought to bear...”
Stuart Gordon of Compact Services (Solutions) Ltd agrees:
“In order to ensure that the images are viewable for posterity we must have the primary archive images in a universal lossless format, currently TIFF however as standards evolve this will inevitably change.

The camera RAW formats are proprietary and are therefore not suitable as the primary archival format; however, if space is available then do archive the RAW as well.

As with all questions it depends on what your goals are and in our case it is ensuring that the archival images are stored in a format that will ensure they are accessible for posterity.”
Paul Wheatley, Digital Preservation Manager at the British Library commented:
“As well as the extremes of the camera/manufacturer specific formats and then at the other end of the spectrum, formats like JPEG2000 and TIFF, we also have a number of emerging RAW file “standards”, of which Open RAW is one. I’ve not seen any independent technology watch work to assess all these options in a reasoned way and make recommendations. Perhaps this is a candidate for a DPC Tech Watch piece? If so, I would like to see authorship from a panel of experts to ensure a balanced view. Both an independent viewpoint and careful analysis of the objectives behind the use of different file formats were a little lacking in the last DPC report. As Larry is hinting at here, I’m sure there are different answers depending on exactly what you want to achieve from the storage and preservation. These aims need to be carefully captured as well as the recommendations to which they relate.”
[UPDATE: Rob Berentsen, of Deloitte Consulting B.V.] brings us back to preservation concerns:
“The discussion whether to use Raw, Tiff or both is one that depends on what the (future) user demands will be and the current preservation demands are. As you can see, Larry prefers Raw so he can use the image in an optimal manner as photographer/manipulator. James however might have a different view on the importance of certain aspects of images and might have different future demands in mind.

Looking at the OAIS model, the most raw image should leave all roads open to create a TIFF image (or others). This TIFF might be more accessible, and even the non-lossless format JPG might be good enough for some people to have [the] image easily accessible over the internet.

To cope with this, the OAIS's archival information package (AIP) offers room for more then just one 'manifestation'. You can see one (the Raw version) as the preservation format, and others are possible future preservation formats or accessible formats. The accessible formats are commonly the ones end-users can get (DIP-package).”
Marc Fresko of Serco suggested there are two questions here:
“One is the question addressed by respondents so far, namely what to do in principle. The answer is clear: best practice is to keep both original and preservation formats.

The second question is "what to do in a specific instance". It seems to me that this has not yet been addressed fully. This second question arises only if there is a problem (cost) associated with the best practice approach. If this is the case, then you have to look at the specifics of the situation, not just the general principles. Most importantly:
  • what will you lose by going from RAW to TIFF (probably nothing that matters for many archival applications, but you will lose something which would matter in other applications)?
  • over how long do you need to preserve access to the images, compared to the likely lifetime of the software you'll need to access the specific RAW format?
You can easily see that different answers to these two questions will lead you to prefer different formats.”
Marc’s comment brings us back to costs, and there is a further series of emails on this topic which I will include in another blog post.

Colin Neilson of the DCC SCARP Project took the discussion towards science data:
“If you generalise and treat your camera as a "Digital Image Acquisition System" :-) then do you keep the original data from the instrument (camera) or would you prefer only to keep the derived data in some form of standardised output?

I suspect in a scientific context you would want to keep both. The problem is that the steps taken in deriving the data are not standardised or documented and may not be repeatable at a future time e.g. when the particular make of instrument (camera) becomes obsolete. The software & algorithms for processing captured image data to standard format are embedded in the camera and function to adjust the data produced from the proprietary image processor (e.g. Canon's DIGIC chip). Some of the "improvements" produced in image quality may not be welcome in a scientific context (e.g. may have introduced artefacts). Some makes and models of camera offer the option save both the "raw" image and the derived image at the cost of memory and processing time.

You can seem some of the issues developed when looking at open microscopy environments where the issue is image capture with effective data management system for multidimensional images including spatial, temporal, spectral and lifetime dimensions. Slide 13 of this presentation by Kevin Eliceiri sets out the current situation. This is quite an exciting area!

Of course this may all be a bit academic (not to say "over the top") if your main concern is to take the best steps to protect you holiday pictures or personal photography efforts. I have found this article useful in an amateur photography context.”
Finally (for this post), Riccardo Ferrante, who is IT Archivist & Electronic Records Program Director for the Smithsonian Institution Archives agreed with Colin:
“We have scientists here that consider the raw data absolutely critical and whose integrity and authenticity must be preserved (archival package). However, the usefulness of the data comes only when the raw data is filtered. The filtered data (derivation) is also considered an archival package by these scientists since it is reused repeatedly in the future, cited in publications, etc. and therefore needs its authenticity, integrity, and provenance to be preserved. To use OAIS-Reference Model language - the DIP generated from the original AIP (raw data) is itself a derived AIP.

In contrast, we have artists who deal solely in photographs and do not consider the raw image to be an object of record, but more like an ingredient for the object that is finalized after post-processing occurs, i.e. after it is published.

To me, it seems the common point is what the creator considers the object of record and the purpose of the raw data in light of that - this should be a major factor in determining the object of record from a recordkeeping viewpoint.”
I thought this was an extremely useful conversation!

Monday, 28 April 2008

PDF: Preserves Data Forever? Hmm…

It’s always good to see new papers coming out of the DPC. Some fantastic work has been undertaken under the DPC banner over the years, and the organisation has done a great job of raising awareness and contributing to approaches addressing digital preservation. But I was somewhat concerned to read the press release announcing their latest Technology Watch – stating that ‘PDF should be used to preserve information for the future’ and ‘the already popular PDF file format adopted by consumers and business alike is one of the most logical formats to preserve today’s electronic information for tomorrow.

This is a fairly controversial statement to make. Yes, PDF can have its uses in a preservation environment, particularly for capturing the appearance characteristic of a document. New versions of the PDF reader tend to render old files in the same way as old viewers. It has the advantage of being an open standard, despite being proprietary, and conversion tools are freely and widely available. But, it is not a magic bullet and there are several potential shortcomings – for example, PDFs are commonly created by non-Adobe applications which return varying quality or functionality in PDF files; it’s not useful for preserving other types of digital records such as emails, spreadsheets, websites or databases; it’s not great for machine parsing (as Owen pointed out in a previous comment on this blog) and there are several issues with the PDF standard which even led to development of PDF/A – PDF for Archiving.

To be fair, the press release does later say that the report suggests adopting PDF/A as a potential solution to the problem of long term digital preservation. And the report itself also focuses more on ‘electronic documents’ than electronic information per se, a generic ‘catch all’ phrase that includes types of information for which PDF is just not suitable.

So what is the report all about? Well, essentially it’s an introduction to the PDF family – PDF/A, PDF/X, PDF/E, PDF/Healthcare, and PDF/UA, in a fairly lightweight preservation context. There is some discussion of alternative formats – including TIFF, ODF, and the use of XML (particularly with regards to XPS, Microsoft's XML Paper Specification) – and an overview of current PDF standards development activities. It’s good to have such an easily digestible overview of the general PDF/A specs and the PDF family. But what I really missed in the report was an in-depth discussion of the practical issues surrounding use of PDF for preservation. For example, how can you convert standard PDF files to PDF/A? How do you convert onwards from PDF/A into another format? In which contexts may PDF/A be unsuitable, for example, in light ofa specific set of preservation requirements? What if you required external content links to remain functional, as PDF/A does not allow external content references – what would you get instead? Would the file contain accessible information to tell you that a given piece of text previously used to link to an external reference? And what exactly is the definition of external reference here – external to the document, or external to the organisation? Should links to an external document in the same records series actually be preserved with functionality intact, and is it even possible?

Speaking of preservation requirements, it would have been particularly useful if the report included a discussion of preservation requirements for formats – this would have informed any subsequent selection or rejection of a format, especially in the section on ‘technologies’. The final section on recommendations hints at this, but does not go into detail. There are also a few choice statements that simply left me wondering – one that really caught my eye was ‘this file format may be less valuable for archival purposes as it may be considered to be a native file format’ (p17), which seems to discount the value of native formats altogether. Perhaps the benefits of submission of native formats alongside PDF representations is a subject which deserves more discussion, particularly for preserving structural and semantic characteristics.

I wholeheartedly agree with perhaps the most pertinent comment in the press release – that PDF ‘should never be viewed as the Holy Grail. It is merely a tool in the armoury of a well thought out records management policy’ (Adrian Brown, National Archives). PDF can have its uses, and the report has certainly encouraged more debate in the organisations I work with as to what those uses are. Time will tell as to whether the debate will become broader still.