Posts tonen met het label APARSEN. Alle posts tonen
Posts tonen met het label APARSEN. Alle posts tonen

zondag 3 juli 2011

How can we prove that digital preservation systems will deliver? (LIBER 5)

david1 This blog post is about the LIBER2011 workshop with the poorest attendance (14 out of 400 conference participants having a choice between three parallel sessions). Attendance may have been poor, but the subject matter was important and thus I can only conclude that I and others who plead the cause of digital preservation still have a lot of work to do. (Or are the other 386 counting on me blogging about it in sufficient detail ;-)

Why testing?

Over the past 15 years or so we have been building preservation systems and putting our digital collections (or, more precisely, ‘digitally encoded information’) into them. But how do we know that they will deliver? Last month in Tallinn, Michael Seadle called our present systems ‘a leap of faith’ and with Andreas Rauber he pleaded for more testing and more exchanges of testing data (see post).

But what do you test? And how?

That was what David Giaretta’s workshop was about, in the context of the APARSEN project (a major European project with 32 partners) (slides in this post courtesy of David Giaretta).

_DSC7084 Giaretta explaining the four phases of APARSEN: Trust, Sustainability, Usability and Access. Testing is part of the trust package.

‘We need more than migration and emulation.’

The most well-known preservation techniques are migration and emulation. The results are tested on the basis of ‘significant properties’: Is the information an organization regards as essential still there after the object has been changed or, alternatively, in the new computer environment that purports to emulate the old computer?

Giaretta asserts that these techniques are useful for some digital objects – and they have a role to play in determining authenticity -, but the techniques do not work for all objects. APARSEN has developed a three-dimensional model to characterize objects technically to be able to determine what tools can be applied:

david2 

Which leads to these conclusions:

david3

So, we need other techniques in addition to migration and emulation. Especially if we want to our information to be part of the Global Brain of Linked Open Data Herbert van de Sompel spoke about on Wednesday.

First question: who do we preserve for?

Giaretta has developed a very elegant way of describing what we do all this work for: we have ‘unfamiliar’ stuff (rows of ones and zeros) which we must make ‘familiar’ for people to be able to use it. We must do that now, and we must continue to do it in the future. Over time, the job will become more difficult.

Second question: what do they need to use the object?

What ‘familiar’ means, depends on the context, on a central concept from the OAIS reference model, the ‘designated community’, the user group an institution works for and their knowledge bases. If the target audience is a group of five-year-olds, our rendering techniques must be very sophisticated so the five-year old only has to push a button. If the audience is a group of computer specialists, less help will be needed.

In this view, the representation information which is part of the OAIS is included in the AIP (OAIS term for the archival information package that includes both the object itself and all the extra information needed to process and render it) becomes the focus of testing the systems (see OAIS Information Model). Is everything there that the designated community needs to be able to use the information?

david4

The ‘representation information network’

As we saw above, the representation information varies between designated communities. But it will also change over time. A present-day computer will understand the information ‘this is XML’. But in 2080 XML is perhaps an archaic file format, and the rendering information will have to be much more specific in telling the computer how it can render XML so a human (or machine) can use it. And if the manual for the programme happens to be in PDF, it will need to include the same information about PDF. Discipline-specific information must also be included, such as vocabularies and ontologies. And when the information package contains a series of dates one must be able to determine the time zone, summer or winter time, etcetera.

It is a network, to which new information must be added as time goes on:

david5

This network can be tested. Is all the required information being preserved?

In the past month, APARSEN has been doing a series of test audits in Europe in preparation for the ISO16363 standard which is in the making. The tests were also designed to test prospective auditors. The provisional conclusions are as follows:

  • most audited organizations do a good job at preserving the bits;
  • quite a few organizations lack succession plans (what happens to the data when my organization ceases to exist?);
  • quite a few have not defined their designated communities;
  • typically, the representation information networks are insufficient or non-existing.

Giaretta concluded:

david6

2011-06-29 11-06-50 - 114

Plenty of empty chairs … (Photo: Jordi Aguilar)

Here is David’s impressive list of references for those of you who want to know more:

1. CCSDS. (2002), Reference model for an Open Archival Information System (OAIS). Retrieved from: http://public.ccsds.org/publications/archive/650x0b1.pdf

2. OAIS update (at the time of writing under CCSDS review), http://public.ccsds.org/sites/cwe/rids/Lists/CCSDS%206500P11/Attachments/650x0p11.pdf

3. Knight, G., 2008, Framework for the definition of significant properties. Retrieved from http://www.significantproperties.org.uk/documents/wp33-propertiesreport-v1.pdf

4. Wilson, A., 2007, Significant Properties Report. Retrieved from http://www.significantproperties.org.uk/documents/wp22_significant_properties.pdf

5. J. Rothenberg and T. Bikson, 1999, 'Carrying Authentic, Understandable and Usable Digital Records Through Time' report to the Dutch National Archives and Ministry of the Interior. Retrieved from http://www.digitaleduurzaamheid.nl/bibliotheek/docs/final-report_4.pdf

6. M. Hedstrom and C.A. Lee, “Significant properties of digital objects: definitions, applications, implications”, Proceedings of the DLM-Forum 2002. Retrieved from http://ec.europa.eu/transparency/archival_policy/dlm_forum/doc/dlm-proceed2002.pdf

7. Cedars project, http://www.leeds.ac.uk/cedars/

8. Investigating the Significant Properties of Electronic Content over time (InSPECT) http://www.significantproperties.org.uk/

9. The InterPARES project, http://www.interpares.org/

10. Wison, A., 2008, Significant Properties of Digital Objects, presented at “What to preserve? Significant Properties of Digital Objects”. Retrieved from http://www.dpconline.org/docs/events/080407sigpropsWilson.pdf

11. DELOS Digital Preservation Testbed. Retrieved from http://www.ifs.tuwien.ac.at/dp/testbed.html

12. OCLC/RLG Working Group on Preservation Metadata, 2002, Preservation Metadata and the OAIS Information Model, A Metadata Framework to Support the Preservation of Digital Objects. Retrieved from http://www.oclc.org/research/projects/pmwg/pm_framework.pdf

13. Derek Sergeant, 2002, Interpretation of the OAIS Model. Retrieved from http://www.erpanet.org/events/2002/copenhagen/presentations/dmserpanet.ppt

14. CASPAR Access Model, http://www.casparpreserves.eu/Members/cclrc/Deliverables/report-on-oais-access-model/at_download/file especially section 2.

15. Michael Factor, Ealan Henis, Dalit Naor, Simona Rabinovici-Cohen, Petra Reshef, Shahar Ronen, IBM Research Lab in Haifa, Israel and Giovanni Michetti, Maria Guercio, University of Urbino, Authenticity and Provenance in Long Term Digital Preservation: Modelling and Implementation in Preservation Aware Storage, TaPP ’09. First Workshop on the Theory and Practice of Provenance. San Francisco, 23 February 2009, http://www.usenix.org/event/tapp09/tech/full_papers/factor/factor.pdf

16. CASPAR Conceptual Model, http://www.casparpreserves.eu/Members/cclrc/Deliverables/caspar-conceptual-model-phase-1-1/at_download/file

17. Giaretta, D., 2007, The CASPAR Approach to Digital Preservation, The International Journal of Digital Curation, Issue 1, Volume 2, http://www.ijdc.net/index.php/ijdc/article/viewFile/29/18

18. CASPAR – Cultural, Artistic and Scientific knowledge for Preservation, Access and Retrieval. See http://www.casparpreserves.eu

19. Mike Coyne, David Duce, Bob Hopgood, George Mallen, Mike Stapleton. The Significant Properties of Vector Images. JISC report, 27 November 2007. http://www.jisc.ac.uk/media/documents/programmes/preservation/vector_images.pdf

20. Mike Coyne, Mike Stapleton. The Significant Properties of Moving Images. JISC report, 26 March 2008. http://www.jisc.ac.uk/media/documents/programmes/preservation/spmovimages_report.pdf

21. Brian Matthews, Brian McIlwrath, David Giaretta, Esther Conway. The Significant Properties of Software: A Study. JISC report, March 2008 http://www.jisc.ac.uk/media/documents/programmes/preservation/spsoftware_report_redacted.pdf

22. Kevin Ashley, Richard Davis, Ed Pinsent. Significant Properties of E-learning Objects. JISC report, March 2008. http://www.jisc.ac.uk/media/documents/programmes/preservation/spelos_report.pdf

23. PARADIGM project, Workbook on Digital Private Papers. http://www.paradigm.ac.uk/workbook/preservation-strategies/file-properties.html

vrijdag 24 juni 2011

Persistent identifiers: policy and ‘will’ vital ingredients (#kepoid)

The world of internet is changeable and volatile. If we are to secure long-term access to content on the internet we have to find mechanisms to bring order to the seeming chaos. Standards, for instance – although I learned last month in Tallinn that we may be rushing into those (see blog post). Persistent identifiers are another type of building blocks for long-term access to digital objects, because PIDs make sure that we can find the object that is being preserved, even if it is moved from one URL to another. But I learned last week that the persistent identifiers are not as persistent as one might hope for. Another illusion down the drain?

_a1

In front of a famous painting by Rembrandt (The Anatomy Lesson of Dr. Nicholaes Tulp), a working group led by Andrew Treloar (standing, at right) dissects the truth about persistent identifiers and their complex relationship with Linked Open Data.

The setting was a two-day seminar on persistent object identifiers (or POID, thus #kepoid) organized by Knowledge Exchange, the PersID project, SURFfoundation and Data Archiving and Networked Service (DANS) in the Hague (14-15 June). Regrettably, I managed to attend only the second day, but it was enough to make me understand how complicated this business is.

This is how it should work: a (national) organization (national library, scientific organization) assigns a unique identifier to a digital object, a so-called persistent identifier. If the object is moved from one URL (internet location) to another, the PI remains the same and a resolver service links the new URL back to the PID.

Borrowing from Andrew Treloar’s presentation (Australian National Data Service), here are the main complications associated with object identifiers:

  • Granularity: what do you assign a PID to? In FRBR terms: to the work? to the expression? to the manifestation? to the item? Or, I may add, to a chapter? to a paragraph? Perhaps we even need multiple PIDs at multiple levels.
  • How do you assign PIDs to objects that are not static, but that change all the time (e.g., databases)?
  • How trustworthy is the object that is being identified (e.g., short url services)?
  • How to point to something inside the object?
  • Who owns the binding between the PID and the object?

And then there is the problem that there are a number of different PID systems (e.g., URN, DOI, PURL), which are not interoperable (comment by Juha Hakala: ‘It is encouraging that it is quite a long time since someone came up with a new PID system.’). And PID’s do not go well together with Linked Open Data (LOD).

_a2

‘Why is it so hard?’ – notes from Jeroen Rombouts’ computer (3TU.Datacenter)

Both Clifford Lynch and Andrew Treloar concluded that solving the technical problems of the PID challenge is the easiest part of the work to be done. Andrew built a pyramid of key success factors (photo above): at the bottom of the pyramid is a sustainability model, the second layer is about policies, the third is about procedures, and the top layer is about will or the intention of individuals to follow the rules and make the system work.

_a5

A room full of persistent identifiers – at right seminar chair Bas Cordewener (SURFfoundation).

In the end the attendees concluded that building interoperability between the existing PID systems is not a top priority. But getting PIDs to work with Linked Data is. Treloar proposed a 'Den Haag manifesto’ to bring this about:

The Hague Manifesto on persistent identifiers and Linked Open Data (LOD) (draft version)

  1. Make sure PID’s can be referred to HTTP URI’s including content negotiation
  2. Use LOD vocabularies, for schema elements
  3. Identify the minimum common set of schema elements, across identifiers in scholarly communication space.
  4. Use same-as relations to help PID interoperability across PID systems/schema’s
  5. Work with the LOD community on simple policies/procedures to improve persistence of HTTP URI’s.

Treloar will work with anybody who is ‘ready, willing and able’ to develop these principles.

Some other recommendations from the meeting:

  • Do an inventory of different PID systems and make transparent how they work, so that organizations contemplating using PID’s know how to choose a system
  • Find the common ground between the systems and use these to widen awareness of PID problems and systems
  • Organize regular meetings between those who are involved in building PID infrastructures to facilitate alignment.

The work is being continued, within PersID and also within the European APARSEN project.

_a4

woensdag 24 november 2010

Onderzoeksdata (3): het Europese landschap

APA Conferentie 2010 Aan het eind van de Alliance (kort: APA, Alliance for Permanent Access)-conferentie (zie ook vorige blogs) is het tijd om de balans op te maken. Wat zijn we opgeschoten ten opzichte van vorig jaar, en wat zijn de vooruitzichten?

aparsen-logo

 

Nieuwe Europese projecten: APARSEN EN ODE

Het APARSEN Network of Excellence volgens projectleider David Giaretta Onder de APA-paraplu zijn twee nieuwe Europese projecten gelanceerd (de website voor beiden is gloednieuw, en dus nog niet helemaal gevuld, maar daar wordt aan gewerkt). Eerst APARSEN (Alliance for Permanent Access to the Records of Science Excellence Network). In dat project werken maar liefst 30 (!) partners samen om een virtueel kennisnetwerk te vormen rond duurzame toegang tot wetenschappelijke bronnen (publicaties, data). Als we de enthousiaste projectleider David Giaretta mogen geloven, gaat dit project alles aan elkaar knopen wat eerder aan onderzoek is gedaan naar duurzame toegankelijkheid (CASPAR, Planets, SHAMAN, etc.) en het vervolgens toepasbaar maken in alle grote e-infrastructuren in Europa. Er zitten grote partners in, dat is zeker. En zelfs het NCDD-bureau gaat een bescheiden bijdrage leveren aan het ‘outreach’-programma. Zoals gisteren al geblogd zal de veelheid aan nationaliteiten en culturen ongetwijfeld ook tot een Toren van Babel leiden, maar gelukkig is een projectleider van de Europese Commissie ervan overtuigd dat het hier een necessary chaos betreft. Ikzelf zoek nog een beetje naar de focus in al die ambities …
Salvatore Mele van CERN, waar o.a. de grote HADRON deeltjesversneller is gehuisvest. ODE (Opportunities for Data Exchange), geleid door Salvatore Mele van CERN, is een ‘compact’ project dat ‘will gather evidence to support the right investment in a data sharing, re-use and preservation layer in the emerging e-Infrastructure’. Ergo: ervoor zorgen dat op de grote onderzoeksinfrastructuren die in Europa ontstaan ook daadwerkelijk een laag wordt gebouwd die ervoor zorgt dat onderzoeksdata duurzaam bewaard, hergebruikt en gedeeld kunnen worden. Want dat is momenteel nog lang niet overal het geval. Volgens Mele is de eerste taak van ODE: ‘collect success stories, near misses and honourable failures in data sharing, re-use and preservation’. Die ga ik in de gaten houden!
Eefke Smit van STM (li) met David Giaretta, Executive Director van APA, en zijn nieuwe boek 'Advanced Digital Preservation' (Springer, jan. 2011)
Overigens werd tijdens de bijeenkomst duidelijk dat hier inderdaad nog een wereld te winnen valt. Eefke Smit, van STM Publishers, liet resultaten uit het PARSE-Insight project zien, waaruit bleek dat de meeste uitgevers inmiddels wel maatregelen hebben genomen om de duurzame toegankelijkheid van hun publicaties veilig te stellen (vaak via bibliotheken, zoals bijvoorbeeld de Nederlandse KB), maar dat ze met data nog niets doen en dat ook niet van plan zijn. Dus blijft de vraag wie daarvoor verantwoordelijkheid neemt én wie ervoor gaat zorgen dat publicaties en onderliggende onderzoeksdata gekoppeld blijven.

 

De Alliance for Permanent Access zelf

De Alliance for Permanent Access (APA) zelf heeft een poosje, wat zal ik zeggen, een beetje gesukkeld. Dat gebeurt soms als de motor achter een organisatie (de executive director) niet de juiste man op de juiste plaats blijkt te zijn. Daarom was het mooi om te constateren dat de nieuwe Executive Director, David Giaretta, voortvarend van start is gegaan. APA leeft weer, en zit vol ambities en plannen. Op een iets andere manier dan bij de oprichting voorzien, namelijk nu ook als thuis voor projecten als ODE en APARSEN. Naast de taak als lobby en aanspreekpunt De APA Participants Meeting na de conferentie voor de Europese Commissie en haar kaderprogramma’s op het gebied van digitale informatie en e-infrastructuren. De Executive Board van APA wordt versterkt met Peter Doorn van DANS, die namens de NCDD zal optreden.
Europa kan ook wel wat hulp gebruiken, want onlangs de indrukwekkende lijst acroniemen die de revue passeerden, en waarachter belangrijke wetenschappelijke projecten en infrastructuren schuilgaan, noemde Salvatore Mele de bijdrage van de Europa aan een omgeving waar onvoorstelbare hoeveelheden data worden geproduceerd, nog maar een ‘drop in the ocean’. Hij vroeg EC-vertegenwoordiger Liina-Maria Munari of de Commissie niet iets kan doen om een sneeuwbaleffect te veroorzaken. Munari antwoordde politiek correct dat ze zich bewust was van de ‘drop in the ocean’, maar dat ze hoopte dat de technische voorzieningen volwassen zouden worden en genoeg plaats zouden bieden voor al die data. En dat de wetenschappers goed data-management zouden integreren in hun werk.

In Kajaani bouwt CSC in een oude papierfabriek een CO2-neutraal datacentrum 

Groen databeheer

Het Finse CSC (IT Center for Science) dat gastheer was voor de conferentie is het grootste wetenschappelijke datacentrum in Finland. Om ervoor te zorgen dat het databeheer CO2-neutraal wordt, bouwt CSC in Kajaani, in noord-Finland, een nieuw datacentrum in een oude papierfabriek. De koeling wordt verzorgd door het koude Finse klimaat en voor de rest van de energiebehoefte zorgen drie (bestaande) waterkrachtcentrales. Een mooi initiatief, dat ik om de een of andere reden niet helemaal kan rijmen met het feit dat de thermostaat in mijn hotelkamer continu op 23 graden stond, terwijl het buiten min 7 was. Maar misschien eis ik nu te veel.
Meer foto's op picasa.
Vincenzo Beruto, ESA

“Sustainability of earth science data preservation is not guaranteed in Europe” (Vincenzo Beruto, European Space Agency, ESA)