woensdag 19 oktober 2011

Digital Preservation Summit (1): how far have we come?


We have gathered in cold, rainy, windy Hamburg today and tomorrow for the Goportis Digital Preservation Summit (#DPS2011). So the sunny note on which Adam Farquhar of the British Library started off the conference was quite welcome - except that some of us (including the undersigned) saw a few more clouds in the sky than Adam. A matter of the cup being half full or half empty?

Adam Farquhar (left) with Angela Dappert (DPC)
Farquhar said our progress to date is 'pretty encouraging'. 'Digital preservation has become business as usual,' he said, 'for large memory institutions.' Now that I reread my notes, the addition about large memory institutions is probably crucial to Adam's argument, but what stuck in my mind, and also in the mind of Steve Knight from the National Library of New Zealand (both of us presented during the day), is that in our opinion, digital preservation is still quite a long way off from being business as usual for most of the stakeholders - including data producers, funders, and all but the very largest memory institutions.
 
Adam also said that we are doing digital preservation 'at a substantial scale', citing recent BL projects involving the migration of millions of objects - whereas in my community (the Netherlands Coalition for Digital Preservation) I hear much grumbling about the (lack of) scalability of the tools we have at our disposal at present.
 

Fortunately (I mean in terms of agreeing on issues), Adam also saw a number of challenges:
  • changes in digital materials (flash, social media with short urls)
  • content in context - when a publication is commented upon over the years, it changes
  • dynamic content - complex objects such as 3D interactive views of crystals; html5 (which incorporates javascript elements)
  • a lack of skills in memory institutions - which is getting worse because of the budget cuts.
And I agree with all those (and surmise that Steve Knight will too).

At the end of his keynote, Adam Farquhar said two things:
  • Do not wait until we know everything to get it right, but do whatever you can now
  • Within our community, we need to become more honest about what works and what does not. That is the only  path to true learning.
To which I can only say: hear! hear!

There is much more good stuff to report from this conference (including instructive disagreements between presenters), but this time live or even semi-live blogging is difficult because I am presenting and moderating myself - plus: this conference is very well organised and the audience does not get any (boring) time off for blogging. Also, it is a 9 to 6 programme, and your blogger needs time to eat and sleep. So, dear readers, I must ask for a little patience. But I assure you: ALL shall be revealed  ... and in a matter of days even, because then we have the KEEP workshop coming up, and iPRES ...

On a more practical note: powerpoint and laptops are wonderful inventions, but can somebody PLEASE come up with a solution whereby every presenter is visible to the audience?
(Thanks to Natalie Walters for allowing me to use this image)

zaterdag 15 oktober 2011

Is emulation something for you (2)

If you tried to register for the The Hague Keep workshop (26-27 October, see last post) but were told the workshop was sold out, you may want to try again. Twenty more seats have been made available.

zaterdag 8 oktober 2011

Is emulation something for you?

KEEP-Keeping-Emulation-Environments-PortableIn their efforts to come up with catchy acronyms, project managers sometimes think of  wonderfully sounding names that, however, tell you too little about what is really going on. The European KEEP project is a case in point to me: KEEP stands for ‘Keeping Emulation Environments Portable’. Perhaps it is just my slow brain and/or the often fuzzy official project language, but for some reason I kept getting visions of shopping bags and briefcases …

jeffrey2On the occasion of the KEEP road show, which is coming to The Hague on 26-27 October (and to Zagreb, 9-11 November, Rome: 29-30 November), I asked KB colleague and KEEP participant Jeffrey van der Hoeven (at left, emulating KEEP user satisfaction) to explain it to me. Here is my version of what he told me:

The most well-known method to deal with software and hardware obsolescence is migration: you change the bits and the bytes of a digital object to make them work on a new platform. However, migration turns out to be not at all as risk-free as we would hope. Plus: it does not work for complex objects such as video games, websites, etc.. An alternative is emulation: you do not change the bits, but write software to make a new computer function as if it were an (old) computer. This means writing emulators for every possible combination (which is a lot of expensive R&D work), but if it works, there are fewer risks involved than with migration.

However, working with emulators is not for dummies. It is technically challenging work for specialists. The KEEP project developed an ‘emulation framework’ that takes care of that. It automatically selects the right emulator and configures the software required to render the object. That sounds quite handy.

Now, what about the ‘portability’? Emulators themselves are pieces of software that become obsolete over time. Therefore, KEEP is developing a KEEP ‘virtual machine’ – that will allow for execution of any software on any platform at any time.

Does this sound too good to be true? Come to The Hague (or Zagreb or Rome) and find out for yourself. Be sure to bring some old obsolete floppies with you to test the systems hands-on.

woensdag 21 september 2011

'The Really Foolproof Solution for Digital Preservation ....

... is ... money ... enough of it, and for an indefinite period. If this cannot be guaranteed, the rest of this book will be essential for you', writes David Giaretta (Alliance for Permanent Access/STFC) at the beginning of his new book. I am not giving you the full title yet, because I am afraid that it might scare you off. Certainly, if I had just seen the title without knowing the book or the background, I would have thought that this book is not for me. And that would have been a mistake, because, although I am not a technical expert, I am thoroughly enjoying it and learning a lot in the process. And so may you, especially if you are one of the many readers that read my post about Giaretta's workshop at the LIBER conference in June. This book gives you much more of where that came from.

David Giaretta with his book.
Contrary to what the title would have you expect (allright then, here it is: Advanced Digital Preservation), there is lots of good solid basic digital preservation information in this book. OAIS for example. Everybody has seen the functional model's diagramme, but how many of us have actually read the standard and understood the philosophy behind it? Giaretta guides us through it in detail. And in pleasantly understandable language. Plus: the importance attached to a clear definition of the designated community in the OAIS model - this is crucial to what follows.

Inevitably, things do get technical in the course of the book; after all, if we did not have technical problems we would not have a digital preservation problem, but the not-too-technical-reader is always warned in good time that this perhaps is a section or a chapter to be skipped. Yet the essence of Giaretta's theory is worth noting for everybody. In his view, migration and emulation, our most well-known preservation strategies, are perhaps good enough for simple objects (PDFs, tiffs, jpegs), but are inadequate for many complex objects which can be found among research data, Giaretta's main focus (hence the title Advanced Digital Preservation). No-one will doubt that scientific research often generates very difficult objects to preserve - they are complex, dynamic, often non-renderable, and so forth.

If you do not preserve research data, this book is still important for you, because other sectors (cultural heritage, archives) that started out with simple objects will increasingly be faced with more complex varieties, as content producers are discovering the extra possibilities and putting them to good use.


The 'Droste' effect

 To tackle the problems of more complex objects, Giaretta, and the CASPAR project team, developed a theory around the Representation Information Network. Simply put: a (or rather: any) data object is nothing but ones and zeros; they must be accompanied by representation information in the metadata to tell you what you need to 'independently interpret, understand and use' (in OAIS language) the data object. The data object can be a single file or multiple files, and the representation information can be anything from a scribbled handwritten note to a complex machine readable formal description (pp 17 ff). In Giaretta's more accessible advocacy language: you have something that is unfamiliar (ones and zeros) and the representation information gives you what you need to make it familiar. However, representation information is not a straight-forward thing: it is more like a set of Russian babushka dolls (in Dutch we would refer to the 'Droste effect', after the cacao nurse that serves from a cacao tin that has her own image on it which serves from a cacao tin that ...): a Word document cannot be understood with Microsoft Office software alone, you will need the operating system, and the programming language, and so forth and so forth. You will need every dictionary, every definition, every standard, every specification that is used somewhere along the line - until you connect with the knowledge base of your designated community, that is: you make the connection with what your designated community has at its disposal in terms of software, hardware and knowledge to work with those.

Over time, as technology evolves, the 'unfamiliarity' of a digital object will increase and the the amount of representation information needed to connect with your designated community will increase with it. Our job is to manage that process and make sure there is always enough representation information to connect with our users. Preferably in an automated way, because there is no way we can do this manually (unless of course we have an truly endless flow of money ...).

Giaretta and his CASPAR team argue that this is the only method that will work for all digital objects, no matter how simple or complicated. The trick will of course be to build that automated process that will keep our digital objects "fresh".

More research is needed to turn this theory into something practical. Meanwhile there is this book to enjoy and learn from, including excursions into non-technical territory: repository audits, preservation chains, business models, stakeholders analysis, and more. Giaretta's fluid style of writing, the many cross-references, summaries, and warning signs have enabled me to delve deeper into the technical level than I thought possible. And I am still learning.

What I would like to see next, however, is more interaction between what Giaretta is developing and what the Open Planets Foundation led by Bram van der Werf (and the related SCAPE project) is working on. What would be really great to have for the community is their joint views on what works and what does not - and in which circumstances, and the direction R&D should take. How about it, gentlemen?

David Giaretta [et al.], Advanced Digital Preservation (Springer, 2011, isbn 978-3-642-16808-6, €99.95).

woensdag 24 augustus 2011

Linked data (3): antwoord op lezersvragen


Het is hoog tijd dat deze blog weer uit zijn zomerslaap ontwaakt, en dat doen we met een thema dat blijkens de vragen die binnenkomen actueel blijft: linked data. Eerder publiceerden we Irene Haslingers inleiding in het thema (Linked Data: wat is dat nu eigenlijk precies?). Daarop kwamen vragen, o.a. van Genoveva Leppaart. René van der Ark van de KB, die onderzoek doet op dit gebied, geeft antwoord:
lod-datasets_2010-09-22_colored
Overzicht van datasets die al opengesteld en gelinkt zijn – stand september 2010 (bron: http://wiki.dbpedia.org/About)

[NB: De software van blogger weigert dienst als we punthaken gebruiken. In onderstaande voorbeelden hebben we die vervangen door [punthaak open] en [punthaak sluit]. Minder fraai, maar dan komt de tekst tenminste op je scherm.]

Voor ik [=René] de onderstaande vragen probeer te beantwoorden moet ik eerst in zijn algemeenheid zeggen dat de eerste stap voor het bereiken van linked data bestaat uit het openstellen van bestaande data voor de buitenwereld via het internet. Het daadwerkelijke linken kan plaats gaan vinden wanneer genoeg partijen hun data open hebben gesteld (bij voorkeur in een veelgebruikt formaat zoals RDF/XML. Dit kan plaatsvinden met een simpele conversie en hoeft dus geenszins de brondata aan te tasten). Hier zit zo onderhand schot in, maar op alle vier de vragen hieronder is nog geen eenduidig antwoord te geven, omdat de discussie hierover nog volop gevoerd wordt.

1. Is het mogelijk Linked Data ook te gebruiken voor de inhoudelijke ontsluiting van bijv. artikelen uit vaktijdschriften, rapporten, boeken, krantenartikelen e.d.? Zo ja, kan je uitleggen hoe dat werkt (op hoofdlijnen)

Ja. Je kunt een open vocabulaire/thesaurus op het web gebruiken om bronnen mee te ontsluiten/verrijken. Het meest bekende voorbeeld, prominent aanwezig in de linked data ‘cloud’ is DBpedia (de wikipedia in RDF/XML formaat), dus deze zal ik voor het gemak als voorbeeld gebruiken. De meest eenvoudige manier om een object te ontsluiten is door de URL van een concept uit (bijvoorbeeld) de DBpedia toe te voegen aan de metadata van het object. Dat zou er dan ongeveer zo uit zien:

[punthaak open]dc:author rdf:resource=”http://dbpedia.org/resource/Albert_Einstein”[punthaak sluit]Albert Einstein[punthaak open]/dc:author[punthaak sluit]

De url waarnaar verwezen wordt in het attribuut blok ‘rdf:resource’ is een verwijzing naar een open data-bron waarmee je effectief ‘linked data’ hebt gecreëerd.

Het idee is dat wanneer je vervolgens de metadata van dit object als open data op het web beschikbaar stelt, andere partijen jouw object geschreven door Albert Einstein kunnen vinden in jouw data, omdat je het hebt ontsloten met het concept Albert Einstein van DBPedia. Dat werkt natuurlijk ook andersom.

Ben je al in het bezit van een eigen thesaurus waarmee objecten ontsloten zijn, dan kan het de investering waard zijn om deze thesaurus te ‘mappen’ (of ‘alignen’) met andere open data-bronnen zoals DBpedia, of, voor persoonsnamen VIAF, of, voor plaatsnamen (wereldwijd), http://geonames.org.

Dit betekent dat je een indirecte link hebt gemaakt (object - ontsloten met dc:author 123456 -  geefMeDeMappingVan(123456) - dbpedia:Albert_Einstein).

2. Het nadeel lijkt me dat je alleen relaties kan terugvinden die vooraf gemaakt (en dus bedacht) zijn. Hoe bepaal je tevoren wat de gebruiker zal willen weten en dus welke relaties je legt? Hoe ver ga je met het leggen van relaties?

Dit hangt heel erg af van de aard en bruikbaarheid van de relatie. In essentie staat elke ontsluitingsterm al in relatie met een object. Dit drukken we in de semantic web/linked data wereld uit als een ‘triple’:

publicatieID dc:author [punthaak open]http://dbpedia.org/resource/Albert_Einstein[punthaak sluit]

Bovenstaand voorbeeld is in de informatiesector in ieder geval een bruikbare relatie. Wanneer het echter om een relatie tussen concepten gaat, wordt het een stuk lastiger:

[punthaak open]http://dbpedia.org/resource/Albert_Einstein[punthaak sluit] dbpedia-owl:spouse dbpedia:Mileva_Marić

(Albert Einstein - heeft huwelijkspartner - Mileva Marić)

Maar: van wanneer tot wanneer waren ze getrouwd? Had hij meerdere vrouwen? Omdat de structuur atomair is, kun je dit soort aanvullende informatie alleen met meer ‘triples’ vastleggen.

3. Hoe kun je in een database met linked data gegevens zoeken? Is dit werk voor de (informatie)professional of kan de eindgebruiker dat ook?

Hier benoem je één van de moeilijkste problemen met semantische zoekmachines. De gemiddelde eindgebruiker, van leek tot wetenschapper, gaat niet de moeite nemen om een zoekvraag semantisch uit te splitsen naar een query die de computer begrijpt. Bovenstaand zoekvoorbeeld zou er dan versimpeld zo uit komen te zien:

Select ?pub Where {
          ?pub dc:author ?auth .
?nat skos:broader [punthaak open]http://thesaurus.org/natuurkundigen[punthaak sluit] .
?auth skos:related ?nat .
?auth hasName ‘Albert Einstein’ .
}

Als een eindgebruiker al de vaardigheden bezit om zo’n zoekvraag te formuleren, dan moet de eindgebruiker ook nog genoeg kennis hebben van de inhoud van de database om erin te kunnen zoeken. Er wordt al sinds deze technologie is bedacht, gezocht naar manieren om googleachtige zoekvragen automatisch te vertalen naar een semantische query, maar dit heeft m.i. nog weinig bruikbaars opgeleverd.

Als je echter alleen de dwarsverbanden tussen thesauri gebruikt, kun je met computers wel een hoop voorwerk doen in het uitbreiden van de kennis over een object en het dus beter ontsluiten; zij het met traditionele en niet met semantische zoektechnieken. Linked data levert dus wel degelijk wat op. Een voorbeeld is een archeologische vondst waarvan alleen de plaatsnaam van de vindplaats in de metadata stond. Als die plaatsnaam wordt gekoppeld aan de database van geonames, dan heb je geocoördinaten tot je beschikking en kun je de vindplaats tekenen op google maps - hiervoor heb je geen semantische database nodig.

4. Wie moeten de relaties gaan aanbrengen? Bibliotheken, informatiecentra en archieven, of uitgevers en/of auteurs? Hoe denk je dat dit geregeld gaat worden?

De consensus in de linked data community is dat dit proces een natuurlijk verloop zal krijgen wanneer genoeg partijen hun data openstellen zodat iedereen ermee aan de slag kan. We kunnen niet voorzien in dit stadium welke partijen kwalitatief en/of kwantitatief de beste links zullen gaan opleveren, als het automatisch gebeurt. Hier is een hoop vertrouwen voor nodig en wanneer de kritieke massa dan bereikt zal zijn, is absoluut niet in te schatten. Wel begint duidelijk te worden dat het openstellen van data in andere sectoren interessante nieuwe technologieën kan op leveren: denk aan de brandweer die gegevens van brandveiligheid vrijgeeft, gekoppeld aan de huizenwaarde in kadastergegevens - dit soort applicaties worden nu op grote schaal gemaakt dankzij het bestaan van ‘open data’.

Of dit proces zich überhaupt gaat voordoen in de informatiesector is iets waar we alleen maar naar kunnen gissen, net als of het iets oplevert voor iemand. Maar het is toch een beetje een kwestie van meegaan in de vaart der volkeren in de hoop dat er iets gebeurt.

In de toekomst is het denk ik wel zo dat bibliotheken/informatiecentra/archieven zich kunnen blijven onderscheiden met de kennis van de eigen collectie en door het faciliteren van betere vindbaarheid; het leggen van relaties tussen verschillende collecties is een onderdeel hiervan. Wie de relaties moet gaan leggen en beheren kan ik niet overzien; wel dat alleen mensen betrouwbare relaties kunnen leggen (machines doen het met een betrouwbaarheid van 80%), maar dat er voor die benodigde menskracht vaak geen budget is.

René van der Ark is projectmedewerker Innovatie & Ontwikkeling bij de Koninklijke Bibliotheek, rene.vanderark@kb.nl

zondag 3 juli 2011

How can we prove that digital preservation systems will deliver? (LIBER 5)

david1 This blog post is about the LIBER2011 workshop with the poorest attendance (14 out of 400 conference participants having a choice between three parallel sessions). Attendance may have been poor, but the subject matter was important and thus I can only conclude that I and others who plead the cause of digital preservation still have a lot of work to do. (Or are the other 386 counting on me blogging about it in sufficient detail ;-)

Why testing?

Over the past 15 years or so we have been building preservation systems and putting our digital collections (or, more precisely, ‘digitally encoded information’) into them. But how do we know that they will deliver? Last month in Tallinn, Michael Seadle called our present systems ‘a leap of faith’ and with Andreas Rauber he pleaded for more testing and more exchanges of testing data (see post).

But what do you test? And how?

That was what David Giaretta’s workshop was about, in the context of the APARSEN project (a major European project with 32 partners) (slides in this post courtesy of David Giaretta).

_DSC7084 Giaretta explaining the four phases of APARSEN: Trust, Sustainability, Usability and Access. Testing is part of the trust package.

‘We need more than migration and emulation.’

The most well-known preservation techniques are migration and emulation. The results are tested on the basis of ‘significant properties’: Is the information an organization regards as essential still there after the object has been changed or, alternatively, in the new computer environment that purports to emulate the old computer?

Giaretta asserts that these techniques are useful for some digital objects – and they have a role to play in determining authenticity -, but the techniques do not work for all objects. APARSEN has developed a three-dimensional model to characterize objects technically to be able to determine what tools can be applied:

david2 

Which leads to these conclusions:

david3

So, we need other techniques in addition to migration and emulation. Especially if we want to our information to be part of the Global Brain of Linked Open Data Herbert van de Sompel spoke about on Wednesday.

First question: who do we preserve for?

Giaretta has developed a very elegant way of describing what we do all this work for: we have ‘unfamiliar’ stuff (rows of ones and zeros) which we must make ‘familiar’ for people to be able to use it. We must do that now, and we must continue to do it in the future. Over time, the job will become more difficult.

Second question: what do they need to use the object?

What ‘familiar’ means, depends on the context, on a central concept from the OAIS reference model, the ‘designated community’, the user group an institution works for and their knowledge bases. If the target audience is a group of five-year-olds, our rendering techniques must be very sophisticated so the five-year old only has to push a button. If the audience is a group of computer specialists, less help will be needed.

In this view, the representation information which is part of the OAIS is included in the AIP (OAIS term for the archival information package that includes both the object itself and all the extra information needed to process and render it) becomes the focus of testing the systems (see OAIS Information Model). Is everything there that the designated community needs to be able to use the information?

david4

The ‘representation information network’

As we saw above, the representation information varies between designated communities. But it will also change over time. A present-day computer will understand the information ‘this is XML’. But in 2080 XML is perhaps an archaic file format, and the rendering information will have to be much more specific in telling the computer how it can render XML so a human (or machine) can use it. And if the manual for the programme happens to be in PDF, it will need to include the same information about PDF. Discipline-specific information must also be included, such as vocabularies and ontologies. And when the information package contains a series of dates one must be able to determine the time zone, summer or winter time, etcetera.

It is a network, to which new information must be added as time goes on:

david5

This network can be tested. Is all the required information being preserved?

In the past month, APARSEN has been doing a series of test audits in Europe in preparation for the ISO16363 standard which is in the making. The tests were also designed to test prospective auditors. The provisional conclusions are as follows:

  • most audited organizations do a good job at preserving the bits;
  • quite a few organizations lack succession plans (what happens to the data when my organization ceases to exist?);
  • quite a few have not defined their designated communities;
  • typically, the representation information networks are insufficient or non-existing.

Giaretta concluded:

david6

2011-06-29 11-06-50 - 114

Plenty of empty chairs … (Photo: Jordi Aguilar)

Here is David’s impressive list of references for those of you who want to know more:

1. CCSDS. (2002), Reference model for an Open Archival Information System (OAIS). Retrieved from: http://public.ccsds.org/publications/archive/650x0b1.pdf

2. OAIS update (at the time of writing under CCSDS review), http://public.ccsds.org/sites/cwe/rids/Lists/CCSDS%206500P11/Attachments/650x0p11.pdf

3. Knight, G., 2008, Framework for the definition of significant properties. Retrieved from http://www.significantproperties.org.uk/documents/wp33-propertiesreport-v1.pdf

4. Wilson, A., 2007, Significant Properties Report. Retrieved from http://www.significantproperties.org.uk/documents/wp22_significant_properties.pdf

5. J. Rothenberg and T. Bikson, 1999, 'Carrying Authentic, Understandable and Usable Digital Records Through Time' report to the Dutch National Archives and Ministry of the Interior. Retrieved from http://www.digitaleduurzaamheid.nl/bibliotheek/docs/final-report_4.pdf

6. M. Hedstrom and C.A. Lee, “Significant properties of digital objects: definitions, applications, implications”, Proceedings of the DLM-Forum 2002. Retrieved from http://ec.europa.eu/transparency/archival_policy/dlm_forum/doc/dlm-proceed2002.pdf

7. Cedars project, http://www.leeds.ac.uk/cedars/

8. Investigating the Significant Properties of Electronic Content over time (InSPECT) http://www.significantproperties.org.uk/

9. The InterPARES project, http://www.interpares.org/

10. Wison, A., 2008, Significant Properties of Digital Objects, presented at “What to preserve? Significant Properties of Digital Objects”. Retrieved from http://www.dpconline.org/docs/events/080407sigpropsWilson.pdf

11. DELOS Digital Preservation Testbed. Retrieved from http://www.ifs.tuwien.ac.at/dp/testbed.html

12. OCLC/RLG Working Group on Preservation Metadata, 2002, Preservation Metadata and the OAIS Information Model, A Metadata Framework to Support the Preservation of Digital Objects. Retrieved from http://www.oclc.org/research/projects/pmwg/pm_framework.pdf

13. Derek Sergeant, 2002, Interpretation of the OAIS Model. Retrieved from http://www.erpanet.org/events/2002/copenhagen/presentations/dmserpanet.ppt

14. CASPAR Access Model, http://www.casparpreserves.eu/Members/cclrc/Deliverables/report-on-oais-access-model/at_download/file especially section 2.

15. Michael Factor, Ealan Henis, Dalit Naor, Simona Rabinovici-Cohen, Petra Reshef, Shahar Ronen, IBM Research Lab in Haifa, Israel and Giovanni Michetti, Maria Guercio, University of Urbino, Authenticity and Provenance in Long Term Digital Preservation: Modelling and Implementation in Preservation Aware Storage, TaPP ’09. First Workshop on the Theory and Practice of Provenance. San Francisco, 23 February 2009, http://www.usenix.org/event/tapp09/tech/full_papers/factor/factor.pdf

16. CASPAR Conceptual Model, http://www.casparpreserves.eu/Members/cclrc/Deliverables/caspar-conceptual-model-phase-1-1/at_download/file

17. Giaretta, D., 2007, The CASPAR Approach to Digital Preservation, The International Journal of Digital Curation, Issue 1, Volume 2, http://www.ijdc.net/index.php/ijdc/article/viewFile/29/18

18. CASPAR – Cultural, Artistic and Scientific knowledge for Preservation, Access and Retrieval. See http://www.casparpreserves.eu

19. Mike Coyne, David Duce, Bob Hopgood, George Mallen, Mike Stapleton. The Significant Properties of Vector Images. JISC report, 27 November 2007. http://www.jisc.ac.uk/media/documents/programmes/preservation/vector_images.pdf

20. Mike Coyne, Mike Stapleton. The Significant Properties of Moving Images. JISC report, 26 March 2008. http://www.jisc.ac.uk/media/documents/programmes/preservation/spmovimages_report.pdf

21. Brian Matthews, Brian McIlwrath, David Giaretta, Esther Conway. The Significant Properties of Software: A Study. JISC report, March 2008 http://www.jisc.ac.uk/media/documents/programmes/preservation/spsoftware_report_redacted.pdf

22. Kevin Ashley, Richard Davis, Ed Pinsent. Significant Properties of E-learning Objects. JISC report, March 2008. http://www.jisc.ac.uk/media/documents/programmes/preservation/spelos_report.pdf

23. PARADIGM project, Workbook on Digital Private Papers. http://www.paradigm.ac.uk/workbook/preservation-strategies/file-properties.html

vrijdag 1 juli 2011

Setting priorities: value shift from printed to digital … and vice versa (LIBER 4)

‘Digital and print is an and and proposition for libraries’, said Graham Jefcoate during the special collections session. But the budgets have not increased. So how should we allocate resources? How should we prioritize?

_DSC6956The Dutch KB has been working on a model that takes an integral view at printed and digital collections. The model can help decide what to spend our money on. At LIBER it was presented by Sophie Ham (photo left) and Tanja de Boer. As it has not been published yet, I gladly offer it here with some detail, because I think it can really help libraries make tough decisions (slides courtesy of Sophie Ham).

The model rates the value of (parts of) collections to determine which ones should be conserved or preserved with priority. Here are the rating criteria:

Primary criteria:

image

Secondary criteria:

image 

And this is how the process works:

image

Multiplying primary and secondary criteria might look like this (just an example):

image

Sophie gave two extreme examples of how such a value assessment can turn out. This example concerns printed versions. Obviously, newspapers score higher on informational value and medieval manuscripts score higher on uniqueness and historic value.

Newspapers

Medieval manuscripts

Informational value

8

5

Aesthetic value

2

9

Historic value

3

9

Use

8

3

Uniqueness

4

10

Condition

2

8

TOTAL

27

44

What happens when these collections are digitized? Some values are transferred to the digital copy (e.g., much of the use and informational value for newspapers), but other values cannot be transferred (e.g., the uniqueness of a medieval manuscript). As Claudia Fabian of the Bayerische Staatsbibliothek demonstrated, the use value of a physical object may even increase when it is digitized, because more people become aware of its existence and become interested.

value1

And here is the end result of this (extreme) example:

 value3

As you can see, the value of the physical medieval manuscript has remained unchanged, whereas some of the value of the printed newspaper collection has been transferred to the digital copy, especially the use value; the KB’s online newspaper database is in great demand by users of all kinds in the Netherlands.

_DSC6951 Discussing the KB model, from the left: Sophie Ham, Claudia Fabian and workshop chair Graham Jefcoate.

So, if push comes to shove, and painful decisions have to be made about whether to build new stacks for physical newspapers or invest in expanding the e-Depot that holds the digital newspapers, this (limited) analysis clearly points in the direction of investing in the e-Depot and perhaps deciding to keep only representative selections of printed newspapers.

Obviously, lots of questions remain. At what granularity should one assess collections? How can one make the assessments as objective as possible?, etc. But it is a promising beginning. If you want to know more or contribute to developing the model, please get in touch with sophie.ham@kb.nl.

_DSC6838 KB colleagues at Sophie’s presentation: from the left, Els van Eijck van Heslinga, Lotte Wilms, Lieke Ploeger,  Victor-Jan Vos.

_DSC7014

Value shift: cool conference bag being put to alternative use. Those with short legs thank the sponsors for the abundance of printed promotional material.