Posts tonen met het label Memento. Alle posts tonen
Posts tonen met het label Memento. Alle posts tonen

donderdag 30 juni 2011

Global Web of Data depends on machine-actionable XML (LIBER 3)

Two inspired keynotes today about the vast new possibilities that machine-readable data – or more precisely: data that machines can act on - open up for the advancement of science. There is so much (digital) data out there that no human can comprehend it all. Fortunately, we have or are developing tireless machines that can publish, merge, search, reason, predict and integrate information. They can establish relationships between fields of science which have never even contemplated getting together and make way for new cross-disciplinary work.

_DSC7172

Herbert van de Sompel, at right, with yesterday’s keynote speaker Rick Luce.

It was, of course, Herbert van de Sompel of Los Alamos who treated the audience of 400 research librarians to a peek into this fascinating world of emerging research possibilities (slides available from slideshare). All based on some rather basic building blocks:

liber2011hvdsompwithblanks-110703101710-phpapp02_Pagina_07 

Enthousiastically Van de Sompel reviewed some of the projects to make all of this possible, starting with his own OAI Object Reuse and Exchange, Open Annotation, and Memento, and then on to other developments, such as the ‘nano publication’, the smallest entity of information that can be searched, merged, read, etc. etc. by machines – a somewhat extended version of an RDF triple. And what to think of Executable Papers – articles that include the software and underlying data so that the reader can repeat the original experiments and draw his own conclusions.

liber2011hvdsompwithblanks-110703101710-phpapp02_Pagina_21

 

Mind boggling! Van de Sompel explained that this is what we need to make it all happen, ‘and it irks me that we have that at our fingertips’:

  • open access to all data
  • permissive (i.e., non-restrictive) creative commons licences
  • money to pay for these tools
  • persistence in identifying the objects – and that is still a challenge (see last week’s post)

From a digital preservation viewpoint, however, there is a complication. As Alma Swan of Enabling Open Scholarship explained a few hours later, it does not work with PDF. PDF was developed for humans to read. Machines cannot read PDF, they need XML.

pdf1

Alma showed that universities are building digital repositories at great speed – there are now almost 2000 of them. But what are we filling them with? Mainly PDF’s … Because they are nice and robust from a preservation viewpoint. And humans can read them.

pdf2

 Alma Swan (seated) awaiting her turn to speak with session chair Bas Savenije.

There is good news as well. As Alma pointed out, digital repositories are attracting new users, mostly users from outside the university who do not have access to licensed digital content from major publishers. Companies, for instance, and private citizens who are making use of the new digital possibilities and getting involved in scientific efforts, such as these:

pdf3

However, much of the material these new users are interested in, is not available in open access, and thus cannot be used by either humans or machines.

So what are we preserving all of the stuff for?

I’m going to sleep on that one …

vip Yours truly will spare no effort to tell you everything; here is my paparazzi shot at the VIP room: LIBER President Paul Ayris (left) and Executive Director Wouter Schallier.

zondag 15 mei 2011

‘Memento’ sparks optimism at closing of IIPC 2011

The IIPC 2011 General Assembly (see previous blog post) came to a close last Friday in an atmosphere of optimism. Martha Anderson of the Library of Congress told me: ‘We did not realize it fully until this conference, but we are very close now to opening up and joining our collections. That is something we have always wanted to do within the IIPC, but we never knew how.’

mementologoAaron Binns of the Internet Archive expressed ‘personal excitement’ when he talked about the Memento project at Los Alamos, which won the DPC Digital Preservation Award last December (see also earlier Dutch blog). ‘Up till now we were still very much working within our own silos. Now for the first time there is a potential framework to work with. Memento can help us realize shared access, which is the whole point of the IIPC.’ Rob Sanderson (aka Azaroth42) of Memento attended the IIPC meeting to report on the progress that is being made.

Helen HockxBinns also reported progress in other areas. Archiving social media and interactive websites remains quite a challenge for the community, he said, but ‘we are making progress. Facebook may be crawlable in the end.’

Another important topic at IIPC 2011 was quality assurance. Checking the quality of harvested websites is presently a labour-intensive manual affair. But Helen Hockx reported on promising work at the British Library to automate quality checks, and Binns expressed confidence that the IIPC community can build on those – although there are still a lot of open questions.

During this General Assembly the IIPC also reached out to its users, as reported yesterday. In retrospect I wonder if the discussions would have gone in another direction if the IIPC had invited humanities researchers in addition to social scientists. Are humanities scholars content with the contents of the web archives? But perhaps it is much too early to assess the value of web archives for researchers in general. As Sophie Ham of the KB, national library of the Netherlands, said during her brief presentation on Monday: ‘If we plant our seeds within the right atmosphere, our diamonds will grow bigger and bigger.’ The KB started harvesting in 2007, so its on-site access to 700 websites is but a beginning of bigger things to come.

During the meeting, the IIPC and the Dutch KB said their farewells to Hilde van Wijngaarden, who in the past ten years developed into one of the driving forces behind digital preservation research & development at the KB and who was an active member of the IIPC steering committee. She is moving on to head the library at the Amsterdam University of Applied Sciences. KB and IIPC will miss her.

Next year’s IIPC General Meeting will be held at the Library of Congress in Washington, DC, the date is yet to be announced, but it is probably going to be May again.

swiss cheese

‘Researchers are used to searching in incomplete collections. For them, I guess, that is part of the fun.’

(Barbara Signori of the Swiss National Library comparing their web archiving efforts to a Swiss cheese)

donderdag 2 december 2010

Memento: een echt geheugen voor Internet

Herbert van de Sompel met een Memento-experiment

Ik kon het gisteravond niet nalaten om een enthousiast YES! de twitteren drie minuten nadat William Kilbride van de Engelse Digital Preservation Coalition had bekendgemaakt dat het project Memento de Digital Preservation Award 2010 had gewonnen. Van alle nominaties was Memento ook mijn favoriet. De ontwikkelaars, Herbert van de Sompel van het Los Alamos National Laboratory en Michael Nelson van Old Dominion University (met een groep collega’s natuurlijk), zijn niet de minsten: zij stonden ook aan de wieg van bijvoorbeeld het Open Archives Initiative (OAI) Metadata Harvesting Protocol (MHP) en Object Reuse and Exchange (ORE). Zij noemen Memento een tijdmachine voor het web – en dat zal ik in een kort mememto-voor-dummies proberen uit te leggen nu ik het verhaal zelf twee keer heb gehoord, één keer in het Vlaams in de KB (hij is Vlaming, en was deze zomer twee maanden lang visiting professor bij DANS), en één keer in het Engels (zie video door Herbert van de Sompel zelf en de technische specificaties, allemaal vrij beschikbaar).

Memento voor dummies

Als je een www-adres (URI) aanroept, krijg je altijd de huidige versie. Wat aan de huidige versie voorafging, is vaak overschreven of anderszins verloren gegaan. Maar voor onderzoekers kan het heel belangrijk zijn om een oude versie te kunnen oproepen. Her en der wordt aan webarchivering gedaan, maar hoe kom je als gebruiker te weten wat er is en waar je het kunt vinden? En als je je oude websites eenmaal hebt gevonden en ze bevatten een link, hoe kom je dan bij de toenmalige versie van die link en niet bij de huidige?

Als je browser met het http-protocol iets gaat opzoeken, zitten in de zoekopdracht al een aantal voorkeuren verstopt, bijvoorbeeld een voorkeurstaal, of een voorkeur voor html-pagina’s in plaats van langzame PDF’s (‘connegs’ of content negotiations). In Memento wordt een ongebruikt deel van die voorkeursinstellingen gebruikt om een datum en tijdstip mee te geven aan de zoekopdracht. Een kind kan de was doen!

Aan de kant van de server waar de website op draait kunnen twee situaties ontstaan: ofwel de server heeft zelf een archief met oude versies (bijvoorbeeld Wikipedia) of de server heeft geen eigen archief. In het eerste geval is toegang vrij gemakkelijk te realiseren. De zoekopdracht wordt naar het archief gestuurd en de versie die het dichtst bij het gevraagde tijdstip ligt wordt weergegeven. De tweede situatie is een stuk gecompliceerder. Want waar bevindt zich het archief of bevinden zich de archieven? In de Memento-logica wordt van dergelijke websites gevraagd dat ze de verzoeken waarin een tijdsbepaling zit niet in behandeling nemen maar doorsturen naar een zogenaamde ‘TimeGate’. Die kan niet alle bestaande webarchieven doorzoeken, dat zou veel te langzaam worden, maar daar zit een API (application programming interface) die de metadata verzamelt van allerlei beschikbare webarchieven en het verzoek doorstuurt naar het webarchief dat het beste antwoord heeft op de vraag.

Even elegant als briljant

Het systeem is eigenlijk heel simpel, maar dat verraadt juist het meesterschap van mensen als van de Sompel c.s. Het is niet zomaar een project maar een bruikbaar systeem dat wereldwijd kan worden ingezet en een enorme stap vooruit betekent voor de doorzoekbaarheid van Internet door de tijden heen. Een terechte winnaar dus.

http://www.mementoweb.org/guide/quick-intro/Randvoorwaarden

Het systeem kan alleen werken als aan bepaalde randvoorwaarden wordt voldaan. Iemand moet die websites archiveren. Iemand moet de metadata van een aantal webarchieven aggregeren in een TimeGate. En het liefst moet Memento een ISO-standaard worden. Wat het laatste betreft: daarover is men al druk in gesprek, want er is enthousiast gereageerd op Memento. Websites archiveren gebeurt ook steeds vaker. En het aggregeren van de metadata? Dat moeten we nog organiseren, op landelijk niveau of misschien zelfs in Europees verband. Volgend voorjaar gaan we een NCDD-conferentie over webarchivering organiseren. Dat lijkt me een goed moment om daar eens naar te kijken.

Lef en innovatie

Ik moet denken aan de presentatie van John Wood op de Europese Alliance conferentie vorige maand en zijn kritiek op hoe in Europa onderzoeksgelden worden verdeeld. Te weinig lef, te weinig innovatie, vond hij. Pat Manson van de Commissie zei tijdens de  iPRES 2010 ook zoiets, dat ze teleurgesteld was over het innovatieve gehalte van de Europese projecten in het kader van de laatste Call for Proposals. Je kunt je afvragen of het poldermodel-waar-iedereen-zijn-duit-in-het-zakje-mag-doen voor technische innovatie te langzaam is, te bureaucratisch. Ik zeg met opzet voor technische innovatie, want daar is inspiratie en creativiteit nodig. Voor organisatorische zaken zoeken we stabiliteit en draagvlak, dat heeft een andere dynamiek. Memento is misschien niet voor niets een Amerikaanse project, financieel mogelijk gemaakt door de collega’s van het National Digital Information Infrastructure and Preservation Program (NDIIPP).