Posts tonen met het label datacuratie. Alle posts tonen
Posts tonen met het label datacuratie. Alle posts tonen

zondag 25 april 2010

UK Digital Curation Centre in Nederland: het stellen van de vragen is belangrijker dan een perfecte optelsom

DSC_0013 Het Engelse Digital Curation Centre (DCC) ontwikkelt al een aantal jaren instrumenten om duurzame toegankelijkheid van onderzoeksdata te bevorderen. Bekende producten zijn het DCC Lifecycle Model en (in samenwerking met andere organisaties) DRAMBORA en DAF. Daar hebben de meesten van ons wel eens van gehoord, maar wie weet er nu echt het fijne van? Het was dus een uitstekend initiatief van het 3TU. Datacentrum om drie trainers van het DCC naar Nederland te halen. Eerst was er een mini-symposium met wetenschappers. Daar was ik niet bij, maar naar verluidt ging het daar 'meer om de wetenschap dan om datamanagement'. In deze tijdens een theepauze opgetekende observatie klinkt enige teleurstelling door, maar ik zou bijna zeggen: wat wil je? Voor ons mag datamanagement dan een doel zijn, voor de wetenschapper is het meestal maar een middel. En ook de organisatoren van 3TU (Alenka Prinčič en Ellen Verbakel) keken er wat filosofischer tegenaan: het is een begin van de dialoog.

Het tweede evenement was een eendaagse cursus 'Curation 101 lite' voor collega's van het 3TU.Datacentrum. Nu betekent '101' in het Angelsaksisch taalgebruik beginnersniveau, en de term 'lite' doet daar nog een schepje bovenop, maar wie naar Delft kwam voor een luchtige vrijdag kwam bedrogen uit. Het  tempo waarin de trainers de DCC Lifecycle, DRAMBORA en DAF de revue lieten passeren lag erg hoog. (Links vlnr Alenka Krinčič van TU Delft, Sara Higgins, Sara Jones, Joy Davidson en Ellen Verbakel van TU Delft). Dat is ook niet verwonderlijk als je bedenkt dat de cursus eigenlijk een driedaags evenement is. Gelukkig kregen we na afloop een dik 'course pack' mee (voor het bijeenrapen waarvan de trainers zowel hun thee- als lunchpauze offerden) waarin we het allemaal nog eens na kunnen lezen.


Wat zijn het voor modellen die het DCC heeft ontwikkeld? Eerst het Lifecycle model (zie ook de Nederlandse vertaling door Ingmar Koch). In feite is dat niet meer of minder dan een logische opdeling in stappen van alles wat er bij duurzame toegankelijkheid komt kijken. Waar moet je allemaal aan denken? Welke functies moet je benoemen en organiseren? Het is vooral gericht op wetenschappelijke onderzoeksgegevens, en idealiter maakt de producent van de gegevens, de onderzoeker, al gebruik van het model bij het opstellen van zijn datamanagementplan. Zoals gebruikelijk bij dit soort modellen is het een heel pak papier, maar ik moet zeggen: als je stap voor stap wordt meegenomen in de logica, dan gaat het wel meer leven. En zoals Joy zelf diverse keren benadrukte: het gaat niet om het precies invullen van alle oefeningen om 'juiste' uitkomsten te krijgen, maar om jezelf de discipline op te leggen om de juiste vragen te stellen -- voordat je ergens aan begint.

DRAMBORA en DAF zijn later (mede door DCC) ontwikkeld, toen in de praktijk bleek dat het Lifecycle Model geen antwoord gaf op alle vragen. DRAMBORA is een instrument dat alles wat je doet in het teken zet van risicomanagement. Wat mij betreft een hele zinvolle benadering. Hoeveel je investeert laat je afhangen van welke risico's voor jouw organisatie het meest bedreigend zijn. Als je hele dienstverlening in elkaar stort wanneer digitale collecties niet toegankelijk zijn, moet je daar snel mee aan de slag. Als niemand het eigenlijk zal merken wanneer een bepaalde collectie niet meer bereikbaar is, dan .... enfin, vul zelf maar in.

DAF (Data Asset Framework of Data Audit Framework, er is weer eens sprake van een naamsverandering) is een instrument dat je kunt inzetten wanneer je bij onderzoekers gaat inventariseren wat voor informatiestromen ze hebben en hoe het is gesteld met hun datamanagementpraktijken. Op basis daarvan kun je gerichte adviezen formuleren ter ondersteuning van het onderzoeksproces.

Aan het eind van de dag zijn er natuurlijk altijd vragen bij de praktische toepasbaarheid van dit soort instrumenten. Want als je het goed wilt doen, gaat er erg veel tijd in zitten, gaf ook Joy Davidson toe. Een collega van de TU Eindhoven vond het allemaal wel erg theoretisch. 'Wij doen gewoon,' zei hij.

Zelf vond ik de ontmoeting erg inspirerend met Mark van Koningsveld, van  http://www.openearth.eu. Wij van de archieven klagen vaak dat het moeilijk is om onderzoekers te verleiden tot het duurzaam opslaan van hun data. De mensen van Deltares ontwikkelden daar een heel praktische en laagdrempelige stimulans voor: zij ontwierpen een manier om allerhande data uit de waterbouwkunde te koppelen aan Google Earth en daarmee fantastische visualisaties te maken. De onderzoekers zijn zo blij met die visualisaties dat ze hun data ervoor beschikbaar willen stellen. Helaas wordt het initiatief bedreigd door geldgebrek. NWO, kunnen jullie daar niet wat aan doen?



Foto in het midden: de trainers offeren hun lunchpauze op om de deelnemers aan een course pack te kunnen helpen.

zaterdag 18 april 2009

Curating research (1): wrap up / samenvatting

Hieronder de eerste impressies die ik tijdens de Curating Research conferentie van gisteren in Den Haag opdeed en als 'wrap-up' aan het slot presenteerde; voor één keer in het Engels. Zie ook de blog hiervoor.

First impressions as I presented them yesterday to the conference Curating Research (The Hague, 17 April 2009) during the final session.

A long, long time ago … that is: early this morninDSC_0009g, the organisers started us out on a very ambitious agenda, I quote Hans Jansen (opening speaker, KB): ‘At the end of the day you will be able to assess the preconditions for implementing long-term preservation in your own organisation – both in terms of policy, technical infrastructure and organisational development.’

So, I ask you, audience: ARE YOU?

[as the room remains quite silent ...]

Perhaps a few highlights will help you answer that question.

DSC_0014 Eileen Fenton (Portico, a US digital archive) painted a very clear general picture of the digital curation landscape: digital information is exploding; we need to manage that information to safekeep it for future generations. And preservation is only a means to the all important end of access.

How do we do it? Well, permanent access does not happen by accident. It takes work, work we still have many uncertainties about, as this is a new field. One thing Eileen told us right away: one size does not fit all, this game is complex.

Eileen had some notable advice for librarians: befriend selection, and get close to the creators of the content. They make critical decisions when it comes to keeping research information accessible in the long term.

Her last, and I think very important point: ‘do not go at this game alone’. Find yourself trusted partners to work with, nationally or internationally.

DSC_0022 Jeffrey van der Hoeven and Tom Kuipers (KB) presented their PARSE.insight project which surveyed how the LIBER libraries deal with digital preservation. To my mind an important outcome of their survey was the low response rate: out of 400 institutions, only 59 completed the questionnaire. What does this say about the current position of research libraries? They are to some degree aware of the issues, but are hesitant to get involved. Perhaps because the issues are too daunting? Or because others should take on the task?

Dale Peters (Göttingen, DRIVER project) reviewed the many European research projects which are under way to tackle the more technical aspects of digital preservation. Although these do assure research libraries that many technical issues are being dealt with on an international scale, I must admit that the quantity and variety of acronyms in this field sometimes overwhelms me. Fortunately, all of them have websites to which you can refer for more detailed information. And if you cannot find your way, send an e-mail to Dale and she will no doubt help you along.

Dale stressed two important points:

a) that we need to do work on linking all the digital information that is out there to serve our clients. I am sure nobody in the room disagrees with that!

b) Also, Dale mentioned – almost in passing – that of course not every repository must by definition have long-term preservation facilities. She agreed with Eileen Fenton that trusted third-party services are not only an acceptable but often an essential part of the digital preservation equation.

Maria Heijne Maria Heijne (TU Delft Library, 3TU Datacentre) agreed with Hans Jansen that securing long-term access to research data and publications is core business for libraries. In her view, libraries have no choice but to engage in data management. She rhetorically asked her audience: who else could do it? It is libraries that have the experience needed, they just need to give their services a digital twist.

This digital twist – as also stressed by Eileen Fenton – involves working very closely together with the research communities themselves. They all have very distinct workflows and metadata schemes which are also very different from libraries’ traditional schemes, so both sides must do a lot of adapting. Although it is early days yet, I think the 3TU.datacentre is really developing into a best practice of research libraries’ involvement with data curation. 3TU do exciting work in developing an entirely new relationship with the research community to create a win-win-situation for researchers and research libraries: better quality data during the research process, which then flows into the digital archive with very little additional effort. Be sure to have another look at her powerpoint presentation when we publish it on our website for more details. And perhaps we can write an article about it in LIBER Quarterly, Maria?

I was very sorry that I could not be in two places this afternoon. Of course I had fellow rapporteurs in the workshops I could not attend, but there was too little time to integrate their notes here. We will, of course, provide a full account in the next issue of LIBER Quarterly.

Here are my own notes from two of the four workshops (with apologies to the other workshop hosts who no doubt had much to say as well):

National and international roles

This session was led by Keith Jeffery (STFC, UK, and chairman of the Alliance for Permanent Access) and Peter Wittenburg (Max Planck Institute of Linguistics, Nijmegen). They focussed their attention on research itself; what elements of the research life cycle should in fact be preserved, and who is responsible for preserving them? This is a monumental question, especially as the researchers in our group kept stressing how complicated research data are. Only the publication is static, everything else is dynamic and thus difficult to preserve.

Some doubts were raised as to whether libraries are in fact best suited for the job of preserving the manifold elements of the research life cycle. Libraries’ work flows and metadata schemes, it was suggested, are perhaps too ‘library-centric’ to serve the research community properly.

So should perhaps the management of live data, including providing access, be separated from the archiving functions? And, more importantly, should communities themselves take care of curation rather than libraries? Krystyna Marek from the European Commission explained that the e-infrastructure vision of the EU is in fact focussing on the research communities themselves.

I should not forget to mention that Hans Geleijnse of LIBER suggested that we draw up 5 or 10 golden rules of digital curation, to help the community along. UNESCO drew up such guidelines in 1996, but they need modernising and updating. Half the attendees of this workshop volunteered on the spot to help bring this about, which I thought was very impressive.

Problems, preconditions and costs: opportunities and pitfalls

DSC_0048 Neil Beagrie (Charles Beagrie Ltd.) took his cue from David Rosenthal, who recently held a controversial presentation at CNI, saying that our real problems now are not about media and hardware obsolescence, as predicted by Jeff Rothenburg in his famous 1995 article, but rather about scale and cost and intellectual property. ‘Bytes are vulnerable to money supply glitches,’ is a memorable quote, especially in these credit crunch times.

DSC_0061 So, what does digital preservation cost? Marcel Ras shared his experiences with the KB e-Depot which now archives about 13 million journal articles, thereby providing a sound base for archiving the published output of research. Between now and 2012, however, the size of the e-Depot will grow expotentially, as the e-Depot will incorporate digitised masters and websites. Yet the cost is expected to remain more or less stable at 6 million euro’s a year, which includes 14 full-time staff.

What does this say about possible costs for research libraries? Neil Beagrie investigated the costs of preserving research data at higher education institutions in the UK; the report 'Keeping Research Data Safe' is on the JISC website. Notable findings are that preserving research data is much more expensive than preserving publications. Also, as predicted earlier this morning, timing is a crucial factor. Good care at creation saves a lot of money in the long run.

Another finding: scale matters. Start-up costs are high, but adding content to existing infrastructures is relatively cheap. The Archaeological Data Service estimates that overall costs tail off substantially anyway with time and scale. This is important for our thinking about funding models and up-front (endowment) payment.

Neil Beagrie concluded his presentation with the observation that when it comes to defining a policy for digital preservation, many higher education institutions still have a long way to go and the same seems to hold true for research libraries.

As a co-organiser of this workshop I would not dare presume that we have answered all your questions, but I do hope that this day has helped you a little further along this no doubt complicated, but also very exciting road.

(The powerpoint presentations will be published next week at http://www.kb.nl/curatingresearch; a full report will appear in the next issue of LIBER Quarterly at http://liber.library.uu.nl/, the current issue of which is devoted entirely to digital preservation issues.)

Photographs, top to bottom (IA): Hans Jansen, KB; Eileen Fenton, Portico; Jeffrey van der Hoeven, KB (ducking behind him his PARSE teammate Tom Kuipers); Maria Heijne, TU Delft Library; Neil Beagrie, Charles Beagrie Ltd., Marcel Ras, KB.

dinsdag 14 april 2009

Voorstel aan Van Dale: datacuratie

Aanstaande vrijdag organiseert de Europese Associatie van Wetenschappelijke Bibliotheken (LIBER) samen met de KB en de NCDD een workshop onder de titel 'Curating Research'. In de vandalePR rond dat evenement hebben we gemerkt dat die term curation nog lang niet overal bekend is.

Data curation is het geheel aan handelingen dat je moet verrichten om data gedurende de hele levenscyclus authentiek, vindbaar en bruikbaar te houden. Uiteraard is duurzaamheid daar een onderdeel van, maar curation is veel meer. Een dermate belangrijk kernbegrip verdient een goede Nederlandse vertaling, zo concludeerde ik vanmiddag aan de koffie met een collega van het Nationaal Archief. Nu kent het Nederlands wel het begrip curator, wat de lading goed lijkt te dekken, maar als werkwoord kent het alleen conserveren en dat lijkt weer net iets te krap voor wat wij proberen te doen met data, namelijk méér dan alleen maar instandhouden.

Dus stel ik voor het begrip datacuratie te gaan gebruiken, lekker aan elkaar zoals wij in het Nederlands plegen te doen. En ja, 'data' is zo ingeburgerd dat ik dat maar zo wil laten.

Als er overwegende bezwaren zijn tegen deze term, dan hoor ik ze graag via het reactieformulier - zo niet, dan zal het eindrapport van de Verkenning het begrip later dit jaar bij Van Dale introduceren.

zondag 1 maart 2009

Data-diensten: generiek of specifiek?


Onlangs schreef Chris Rusbridge van het UK Digital Curation Centre een interessante blog over de diverse manieren waarop je infrastructuren - in dit geval voor de wetenschap - kunt insteken: lokaal, nationaal, internationaal, en, misschien nog belangrijker: generiek of discipline-specifiek.

Rusbridge reageerde op recente initiatieven om een nationale infrastructuur voor wetenschappelijke data te bouwen, het UKRDS initiatief. Hij constateert dat de eisen die je aan een infrastructuur voor wetenschappelijke informatie moet stellen eigenlijk haaks op elkaar staan:

- voor datacentra en digitale archieven geldt: hoe groter des te goedkoper, en hoe generieker hoe gemakkelijker ze op te hangen zijn aan bestaande structuren als universiteitsbibliotheken met hun langetermijnfinanciering.
- voor goede datacuratie geldt: hoe lokaler hoe effectiever, omdat iedere discipline weer zijn eigen eisen stelt aan hoe data georganiseerd moet worden. Die curatie zou je dus het liefst zo dicht mogelijk bij de wetenschapsuitoefening willen onderbrengen, in onderzoeksprojecten die vaak tijdelijk van aard zijn.

Hoe breng je die twee uitersten bij elkaar? Rusbridge suggereert een soort matrix met generieke archieven, waarbij de instroom van wetenschappers zelf komt (eigen verantwoordelijkheid!), die dan weer moeten voldoen aan curatienormen die nationaal bepaald worden door koepels, en door de financiers verplicht worden gesteld. En daar past hij bestaande instellingen als JISC en het Research Information Network dan weer in als adviseurs.

Rusbridge signaleert hier een belangrijk knelpunt: economie en wetenschap stellen verschillende eisen. Dat kan niet anders dan leiden tot een complexe infrastructuur, waarbij het heel belangrijk is om rollen en verantwoordelijkheden heel goed te benoemen. Complicatie daarbij is dat er maar weinig stakeholders zijn die het geheel overzien, dat blijkt ook uit de interviews in het kader van de Nationale Verkenning.

Wat Nederland betreft zie ik nog niet direct de impuls om een 'nationale' wetenschappelijke data-service te gaan bouwen; misschien zijn we daarin als klein land anders dan de UK. Maar de uiteenlopende eisen van economie en wetenschap gelden ook hier.