{"id":2810,"date":"2014-04-17T21:11:37","date_gmt":"2014-04-17T21:11:37","guid":{"rendered":"https:\/\/digitisation.eu\/blog\/?p=1747"},"modified":"2014-04-17T21:11:37","modified_gmt":"2014-04-17T21:11:37","slug":"how-to-maximise-usage-of-digital-collections","status":"publish","type":"post","link":"https:\/\/digitisation.eu\/?p=2810","title":{"rendered":"How to maximise usage of digital collections"},"content":{"rendered":"<p><span style=\"color: #808080;\"><em>(Reblogged from <a title=\"Research in KB \" href=\"http:\/\/researchkb.wordpress.com\/\" rel=\"home\"><span style=\"color: #808080;\">Research in KB<\/span><\/a>\u00a0blog:\u00a0http:\/\/researchkb.wordpress.com\/2014\/04\/13\/how-to-maximise-usage-of-digital-collections\/)<\/em><\/span><\/p>\n<p>Libraries want to understand the researchers who use their digital collections and researchers want to understand the nature of these collections better. The seminar \u2018Mining digital repositories\u2019 brought them together at the Dutch Koninklijke Bibliotheek (KB) on 10-11 April, 2014,<!--more--> to discuss both the good and the bad of working with digitised collections \u2013 especially newspapers. And to look ahead at\u00a0what a \u2018digital utopia\u2019 might look like. One easy point to agree on: it would be a world with less restrictive copyright laws. And a world where digital \u2018portals\u2019 are transformed into \u2018platforms\u2019 where researchers can freely \u2018tinker\u2019 with the digital data. \u2013\u00a0<em>Report &amp; photographs by Inge Angevaare, KB.<\/em><\/p>\n<figure style=\"width: 640px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" alt=\"Mining Digital Repositories Conference 2014\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_21054_4.jpg?w=640&amp;h=288\" width=\"640\" height=\"289\" \/><figcaption class=\"wp-caption-text\">Hans-Jorg Lieder of the Berlin State Library (front left) is given an especially warm welcome by conference chair Toine Pieters (Utrecht), \u2018because he was the only guy in Germany who would share his data with us in the Biland project.\u2019<\/figcaption><\/figure>\n<h1>Libraries and researchers: a changing relationship<\/h1>\n<p>\u2018A lot has changed in recent years,\u2019 Arjan van Hessen of the University of Twente and the CLARIN project told me. \u2018Ten years ago someone might have suggested that perhaps we should talk to the KB. Now we are practically in bed together.\u2019<\/p>\n<p>But each relationship has its difficult moments. Researchers are not happy when they discover gaps in the data on offer, such as missing issues or volumes of newspapers. Or incomprehensible transcriptions of texts because of inadequate OCR (<a href=\"http:\/\/en.wikipedia.org\/wiki\/Optical_character_recognition\">optical character recognition<\/a>). Conference organisers Toine Pieters and Jaap Verheul (University of Utrecht) invited\u00a0<strong>Hans-Jorg Lieder<\/strong>\u00a0of the Berlin State Library to explain why he \u2018could not give researchers everything everywhere today\u2019.<\/p>\n<h1>Lieder &amp; Thomas: \u2018Digitising newspapers is difficult\u2019<\/h1>\n<p>Both\u00a0<strong>Deborah Thomas<\/strong>\u00a0of the Library of Congress and\u00a0<strong>Hans-Jorg Lieder<\/strong>\u00a0stressed how complicated it is to digitise historical newspapers. \u2018OCR does not recognise the layout in columns, or the \u201ccontinued on page 5\u2033. Plus the originals are often in a bad state \u2013 brittle and sometimes torn paper, or they are bound in such a way that text is lost in the middle. And there are all these different fonts, e.g., Gothic script in German, and the well-known long-s\/f confusion.\u2019\u00a0Lieder provided the ultimate proof of how difficult digitising newspapers is: \u2018Google only digitises books, they don\u2019t touch newspapers.\u2019<\/p>\n<figure style=\"width: 640px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" \" alt=\"Mining Digital Repositories Damaged Newspapers\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2002whatlibrarieshavetoworkwith.jpg?w=640&amp;h=426\" width=\"640\" height=\"427\" \/><figcaption class=\"wp-caption-text\">Thomas: \u2018The stuff we are digitising is often damaged\u2019<\/figcaption><\/figure>\n<p>Another thing researchers should be aware of:\u00a0\u2018Texts are liquid things. Libraries enrich and annotate texts, versions may differ.\u2019 Libraries do their best to connect and cluster collections of newspapers (e.g., in the\u00a0<a href=\"http:\/\/www.europeana-newspapers.eu\/\">Europeana Newspapers<\/a>), but \u2018the truth of the matter is that most newspapers collections are still analogue; at this moment we have only bits and pieces in digital form, and there is a lot of bad OCR.\u2019 There is no question that libraries are working on improving the situation, but funding is always a problem.\u00a0And the choices to be made with bad OCR are sometimes difficult: Should we manually correct it all, or maybe retype it, or maybe even wait a couple of years for OCR technology to improve?\u2019<\/p>\n<figure style=\"width: 300px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" alt=\"Mining Digital Repositories Conference Claeyssens Van Hessen Kenter\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2145claeyssensvanhessen2.jpg?w=300&amp;h=213\" width=\"300\" height=\"213\" \/><figcaption class=\"wp-caption-text\">Librarians and researchers discuss what is possible and what not. From the left, Steven Claeyssens, KB Data Services, Arjan van Hessen, CLARIN, and Tom Kenter, Translantis.<\/figcaption><\/figure>\n<h1>Researchers: how to mine\u00a0for meaning<\/h1>\n<p>Researchers themselves are debating how they can fit these new digital\u00a0resources into their academic work. Obviously, being able to search millions of newspaper pages from different countries in a matter of days opens up a lot of new research possibilities. Conference organisers\u00a0<strong>Toine Pieters\u00a0<\/strong>and<strong>\u00a0Jaap Verheul<\/strong>\u00a0(University of Utrecht) are both involved in the<a href=\"http:\/\/heranet.info\/welcome-hera-humanities-the-european-research-area\">HERA<\/a>\u00a0<a href=\"http:\/\/translantis.nl\/\">Translantis<\/a>\u00a0project which is taking a break from traditional \u2018national\u2019 historical research by looking at transnational influences of so-called \u2018reference cultures\u2019:<\/p>\n<figure style=\"width: 300px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2053referenceculture.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Mining digital repositories - Definition of reference cultures\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2053referenceculture.jpg?w=300&amp;h=195\" width=\"300\" height=\"195\" \/><\/a><figcaption class=\"wp-caption-text\">Definition of Reference Cultures in the Translantis project which mines digital newspaper collections<\/figcaption><\/figure>\n<p><span style=\"line-height: 1.714285714; font-size: 1rem;\">In the 17th century the Dutch Republic was such a reference culture. In the 20th century the United States developed into a reference culture and Translantis digs deep into the digital newspaper archives of the Netherlands, the UK, Belgium and Germany to try and find out how the United States is depicted in public discourse:<\/span><\/p>\n<figure style=\"width: 640px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2059jaapverheul.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Mining Digital Repositories Jaap Verheul Translantis\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2059jaapverheul.jpg?w=640&amp;h=471\" width=\"640\" height=\"471\" \/><\/a><figcaption class=\"wp-caption-text\">Jaap Verheul (Translantis) shows how the US is depicted in Dutch newspapers<\/figcaption><\/figure>\n<p><strong style=\"line-height: 1.714285714; font-size: 1rem;\">Joris van Eijnatten<\/strong><span style=\"line-height: 1.714285714; font-size: 1rem;\">\u00a0introduced another\u00a0transnational HERA project,\u00a0<\/span><a style=\"line-height: 1.714285714; font-size: 1rem;\" href=\"http:\/\/heranet.info\/asymenc\/index\">ASYMENC<\/a><span style=\"line-height: 1.714285714; font-size: 1rem;\">, which is exploring cultural aspects of European identity with digital humanities methodologies.<\/span><\/p>\n<p>All of this sounds straightforward enough, but researchers\u00a0themselves have yet to develop a scholarly culture around the new resources:<\/p>\n<ul>\n<li>What type of research questions do the digital collections allow? Are these new questions or just old questions to be researched in a new way?<\/li>\n<li>What is scientific \u2018proof\u2019 if the collections you mine have big gaps and faulty OCR?<\/li>\n<li>How to interpret the findings? You can search words and combinations of words in digital repositories, but\u00a0how can you assess what the words\u00a0<em>mean<\/em>? Meanings change over time. Also: how can you distinguish between irony and seriousness?<\/li>\n<li>How do you know that a repository is trustworthy?<\/li>\n<li>How to deal with language barriers in transnational research?\u00a0Mere translations of concepts do not reflect the sentiment behind the words.<\/li>\n<li>How can we analyse what newspapers do\u00a0<em>not<\/em>\u00a0discuss (also known as the \u2018Voldemort\u2019 phenomenon)?<\/li>\n<li>How sustainable is digital content? Long-term storage of digital objects is uncertain and expensive. (Microfilms are much easier to keep, but then again, they do not allow for text mining \u2026)<\/li>\n<li>How do available tools influence research questions?<\/li>\n<li>Researchers need a better understanding of text mining per se.<\/li>\n<\/ul>\n<h1>Some humanities scholars have yet to be convinced of the need to go digital<\/h1>\n<p><strong>Rens Bod<\/strong>, Director of the Dutch\u00a0<a href=\"http:\/\/cdh.uva.nl\/\">Centre for Digital Humanities<\/a>\u00a0enthusiastically presented his ideas about the value of\u00a0<a href=\"http:\/\/en.wikipedia.org\/wiki\/Parsing\">parsing<\/a>\u00a0(analysing parts of speech) for uncovering deep patterns in digital repositories. If you want to know more:\u00a0<a href=\"http:\/\/ukcatalogue.oup.com\/product\/9780199665211.do\">Bod recently published a book<\/a>\u00a0about it.<\/p>\n<figure style=\"width: 238px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2258rensbod.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Rens Bod\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2258rensbod.jpg?w=238&amp;h=300\" width=\"238\" height=\"300\" \/><\/a><figcaption class=\"wp-caption-text\">Professor Rens Bod: \u2018At the University of Amsterdam we offer a free course in working with digital data.\u2019<\/figcaption><\/figure>\n<p><span style=\"line-height: 1.714285714; font-size: 1rem;\">But\u00a0in the context of this blog, his remarks about the lack of big data awareness and\u00a0competencies\u00a0among many humanities scholars, including young students, was perhaps more striking. The University of Amsterdam offers\u00a0<\/span><a style=\"line-height: 1.714285714; font-size: 1rem;\" href=\"http:\/\/cdh.uva.nl\/news-and-events\/events\/content\/workshops\/2013\/10\/digital-tools-in-the-humanities.html\">a crash course in working with digital data<\/a><span style=\"line-height: 1.714285714; font-size: 1rem;\">\u00a0to bridge the gap. The one-week, free course, deals with all aspects of working with data, from \u2018gathering data\u2019 to \u2018cooking data\u2019.<\/span><\/p>\n<p>As the scholarly dimensions of working with big data are not this blogger\u2019s expertise, I will not delve into these further but gladly refer you to an article Toine Pieters and Jaap Verheul are writing about the scholarly outcomes of the conference [I will insert a link when it becomes available].<\/p>\n<figure style=\"width: 300px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2088verheulpieters.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Mining Digital Repositories Jaap Verheul Toine Pieters\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2088verheulpieters.jpg?w=300&amp;h=205\" width=\"300\" height=\"205\" \/><\/a><figcaption class=\"wp-caption-text\">Conference hosts Jaap Verheul (left) and Toine Pieters taking analogue notes for their article on Mining Digital Repositories. And just in case you wonder: the meeting rooms are probably the last rooms in the KB to be migrated to Windows 7<\/figcaption><\/figure>\n<h1>More data providers: the \u2018bad\u2019 guys in the room<\/h1>\n<p>It was the commercial data providers in the room themselves that spoke of \u2018bad guys\u2019 or \u2018bogey man\u2019 \u2013 an image both\u00a0<strong>Ray Abruzzi<\/strong>\u00a0of\u00a0<a href=\"http:\/\/www.cengage.co.uk\/\">Cengage Learning\/Gale<\/a>\u00a0and\u00a0<strong>Elaine Collins<\/strong>\u00a0of<a href=\"http:\/\/www.dcthomsonfamilyhistory.com\/\">DC Thomson Family History<\/a>\u00a0were hoping to at least soften a bit. Both companies provide huge quantities of digitised material. And, yes, they are in it for the money, which would account for their bogeyman image. But, they both stressed, everybody benefits from their efforts:<\/p>\n<figure style=\"width: 640px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2189thomson.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Value proposition of DC Thomson Family History\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2189thomson.jpg?w=640&amp;h=478\" width=\"640\" height=\"479\" \/><\/a><figcaption class=\"wp-caption-text\">Value proposition of DC Thomson Family History<\/figcaption><\/figure>\n<p><span style=\"line-height: 1.714285714; font-size: 1rem;\">Cengage Learning is putting 25-30 million pages online annually. Thomson is digitising 750 million (!) newspaper &amp; periodical pages for the British Library. Collins: \u2018We take the risk, we do all the work, in exchange for certain rights.\u2019 If you want to access the archive, you have to pay.<\/span><\/p>\n<p>In and of itself, this is quite understandable. Public funding just doesn\u2019t cut it when you are talking billions of pages. Both the KB\u2019s Hans Jansen and Rens Bod (U. of Amsterdam) stressed the need for public\/private partnerships in digitisation projects.<\/p>\n<p>And yet.<\/p>\n<p>Elaine Collins readily admitted that researchers \u2018are not our most lucrative stakeholders\u2019; that most of Thomson\u2019s revenue comes from genealogists and the general public. So why not give digital humanities scholars free access to their resources for research purposes, if need be under the strictest conditions that the information does not go anywhere else? Both Abruzzi and Collins admitted that such restricted access is difficult to organise. \u2018And once the data are out there, our entire investment is gone.\u2019<\/p>\n<p><strong>Libraries to mediate access?<\/strong><\/p>\n<p>Perhaps, Ray Abruzzi allowed, access to certain types of data, e.g., metadata, could be allowed\u00a0under certain conditions, but, he stressed, individual scholars who apply to Cengage for access do not stand a chance. Their requests for data are far too varied for Cengage to have any kind of business proposition. And there is the trust issue. Abruzzi recommended that researchers turn to libraries to mediate access to certain content. If libraries give certain guarantees, then perhaps \u2026<\/p>\n<figure style=\"width: 283px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2121pieters.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Mining Digital Repositories Toine Pieters\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2121pieters.jpg?w=283&amp;h=300\" width=\"283\" height=\"300\" \/><\/a><figcaption class=\"wp-caption-text\">You think OCR is difficult to read? Try human handwriting!<\/figcaption><\/figure>\n<h1>What do researchers want from libraries?<\/h1>\n<p>More data, of course, including more contemporary data (\u2026 ah, but copyright \u2026)<\/p>\n<p>And better quality OCR, please.<\/p>\n<p>What if libraries have to choose between quality and quantity? \u00a0That is when things get tricky, because the answer would depend on the researcher you question. Some may choose quantity, others quality.<\/p>\n<p>Should libraries build tools for analysing content? The researchers in the room seemed to agree that libraries should concentrate on data rather than tools. Tools are very temporary, and researchers often need to build the tools around their specific research questions.<\/p>\n<p>But it would be nice if libraries started allowing users to upload enrichments to the content, such as better OCR transcriptions and\/or metadata.<\/p>\n<figure style=\"width: 640px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2009roomthomas.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Mining Digital Repositories 2014\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2009roomthomas.jpg?w=640&amp;h=275\" width=\"640\" height=\"275\" \/><\/a><figcaption class=\"wp-caption-text\">Researchers and libraries discussing what is desirable and what is possible. In the front row, from the left, Irene Haslinger (KB), Julia Noordegraaf (U. of Amsterdam), Toine Pieters (Utrecht), Hans Jansen (KB); further down the front row James Baker (British Library) and Ulrich Tiedau (UCL). Behind the table Jaap Verheul (Utrecht) and Deborah Thomas (Library of Congress).<\/figcaption><\/figure>\n<p><span style=\"line-height: 1.714285714; font-size: 1rem;\">And there is one more urgent request: that libraries become more transparent in what is in their collections and what is not. And be more open about the quality of the OCR in the collections. Take, e.g., the new Dutch national search service\u00a0<\/span><a style=\"line-height: 1.714285714; font-size: 1rem;\" href=\"http:\/\/www.delpher.nl\/\">Delpher<\/a><span style=\"line-height: 1.714285714; font-size: 1rem;\">. A great project, but scholars must know\u00a0<\/span><em style=\"line-height: 1.714285714; font-size: 1rem;\">exactly<\/em><span style=\"line-height: 1.714285714; font-size: 1rem;\">\u00a0what\u2019s in it and what\u2019s not for their findings to have any meaning. And for scientific validity they must be able to reconstruct such information in retrospect. So a full historical overview of what is being added at what time would be a valuable addition to Delpher. (I shall personally communicate this request to the Delpher people, who are, I may add, working very hard to implement user requests).<\/span><\/p>\n<figure style=\"width: 640px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_1960americannewspapers.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"American newspapers\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_1960americannewspapers.jpg?w=640&amp;h=506\" width=\"640\" height=\"506\" \/><\/a><figcaption class=\"wp-caption-text\">Deborah Thomas of the US Library of Congress: \u2018This digital age is a bit like the American Wild West. It is a frontier with lots of opportunities and hopes for striking it rich. And maybe it is a bit unruly.\u2019<\/figcaption><\/figure>\n<h1>New to the library: labs for researchers<\/h1>\n<p>Deborah Thomas of the Library of Congress made no bones about her organisation\u2019s strategy towards researchers: We put out the content, and you do with it whatever you want. In addition to API\u2019s (<a href=\"http:\/\/en.wikipedia.org\/wiki\/Application_programming_interface\">Application Protocol Interfaces<\/a>), the Library is also allowing for downloads of bulk content. The basic content is available free of charge, but additional metadata levels may come at a price.<\/p>\n<p>The British Library (BL) is taking a more active\u00a0approach. The BL\u2019s\u00a0<strong>James Baker<\/strong>\u00a0explained how the BL is trying to bridge the gap between researchers and content by providing special<a href=\"https:\/\/gist.github.com\/drjwbaker\/10355412\">labs for researchers<\/a>. As I (unfortunately!) missed that parallel session, let me mention the KB\u2019s own efforts to set up a\u00a0<a href=\"http:\/\/lab.kbresearch.nl\/scansion.info.html\">KB lab<\/a>\u00a0where researchers are invited to experiment with KB data making use of open source tools. The lab is still in its \u2018pre-beta phase\u2019 as\u00a0<strong>Hildelies Balk<\/strong>\u00a0of the KB explained. If you want the full story, by all means attend the\u00a0<a href=\"http:\/\/dhbenelux.org\/dhbenelux-2014-conference\/\">Digital Humanities Benelux Conference<\/a>\u00a0in the Hague on 12-13 June, where Steven Claeyssens and Clemens Neudecker of the KB are scheduled to launch the beta-version of the platform. Here is a sneak preview of the lab, a scansion machine built by KB Data Services in collaboration with phonologist Marc van Oostendorp (audio in Dutch):<br \/>\n<iframe loading=\"lazy\" src=\"\/\/www.youtube.com\/embed\/FcTufco9P3A\" height=\"390\" width=\"640\" allowfullscreen=\"\" frameborder=\"0\"><\/iframe><\/p>\n<h1>Europeana: the aggregator<\/h1>\n<blockquote><p>\u201cPortals are for visiting; platforms are for building on.\u201d<\/p><\/blockquote>\n<p>Another effort by libraries to facilitate transnational research is the aggregation of their content in Europeana, especially\u00a0<a href=\"http:\/\/www.europeana-newspapers.eu\/\">Europeana Newspapers<\/a>. For the time being the<em>metadata<\/em>\u00a0are being aggregated, but in\u00a0<strong>Alistair Dunning<\/strong>\u2018s vision, Europeana will grow from an end-user portal into a data brain, a cloud platform that will include the content and allow for metadata enrichment:<\/p>\n<figure style=\"width: 300px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2214alistairdunningeuropeana.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Alistair Dunning: 'Europeana must grow into\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2214alistairdunningeuropeana.jpg?w=300&amp;h=212\" width=\"300\" height=\"212\" \/><\/a><figcaption class=\"wp-caption-text\">Alistair Dunning: \u2018Europeana must grow into a data brain to bring disparate data sets together.\u2019<\/figcaption><\/figure>\n<figure style=\"width: 300px\" class=\"wp-caption aligncenter\"><a style=\"line-height: 1.714285714; font-size: 1rem;\" href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2221europeana.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Dunning's vision of Europeana in the future\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_2221europeana.jpg?w=300&amp;h=282\" width=\"300\" height=\"282\" \/><\/a><figcaption class=\"wp-caption-text\">Dunning\u2019s vision of Europeana 3.0<\/figcaption><\/figure>\n<p><span style=\"line-height: 1.714285714; font-size: 1rem;\">Dunning also indicated\u00a0that Europeana might develop brokerage services to clear content for non-commercial purposes. In a recent\u00a0<\/span><a style=\"line-height: 1.714285714; font-size: 1rem;\" href=\"http:\/\/www.europeana-newspapers.eu\/qa-with-newspapers-researchers-toine-pieters\/\">interview Toine Pieters\u00a0<\/a><span style=\"line-height: 1.714285714; font-size: 1rem;\">said\u00a0that researchers would welcome Europeana to take such a role, \u2018because individual researchers should not be bothered with all these access\/copyright issues.\u2019 In the United States, the Library of Congress is not contemplating a move in that direction, Deborah Thomas told her audience. \u2018It is not our mission to negotiate with publishers.\u2019 And recent \u2018Mickey Mouse\u2019 legislation, said to have been inspired by Disney interests, seems to be leading to less rather than more access.<\/span><\/p>\n<h1>Dreaming of digital utopias<\/h1>\n<p>What would a digital utopia look like for the conference attendees? Jaap Verheul invited his guests\u00a0to dream of what they would do if they were granted, say, \u20ac100 million to spend as they pleased.<\/p>\n<p>Deborah Thomas of the Library of Congress would put her money into partnerships with commercial companies to digitise more material, especially the post-1922 stuff (less restrictive copyright laws being part and parcel of the dream). And she would build facilities for uploading enrichments to the data.<\/p>\n<p>James Baker of the British Library would put his money into the labs for researchers.<\/p>\n<p>Researcher Julia Noordegraaf of the University of Amsterdam (heritage and digital culture) would rather put the money towards improving OCR quality.<\/p>\n<p>Joris van Eijnatten\u2019s dream took the Europeana plans a few steps further. His dream would be of a \u2018Globiana 5.0\u2032 \u2013 a worldwide, transnational repository filled with material in standardised formats, connected to bilingual and multilingual dictionaries and researched by a network of multilingual, big data-savvy researchers. In this context, he suggested that \u2018Google-like companies might not be such a bad thing\u2019 in terms of sustainability and standardisation.<\/p>\n<figure style=\"width: 238px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_1974veijnatten.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Joris van Eijnatten\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_1974veijnatten.jpg?w=238&amp;h=300\" width=\"238\" height=\"298\" \/><\/a><figcaption class=\"wp-caption-text\">Joris van Eijnatten: \u2018Perhaps \u2013 and this is a personal observation \u2013 Google-like companies are not such a bad thing after all in terms of sustainability and standardisation of formats.\u2019<\/figcaption><\/figure>\n<p><span style=\"line-height: 1.714285714; font-size: 1rem;\">At the end of the two-day workshop, perhaps not all of the ambitious agenda had been covered. But, then again, nobody had expected that.<\/span><\/p>\n<figure style=\"width: 300px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_1969agenda.jpg\"><img loading=\"lazy\" decoding=\"async\" alt=\"Agenda for Mining Digital Repositories 2014\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/adsc_1969agenda.jpg?w=300&amp;h=180\" width=\"300\" height=\"180\" \/><\/a><figcaption class=\"wp-caption-text\">Mining Digital Repositories 2014 \u2013 the ambitious agenda<\/figcaption><\/figure>\n<p><span style=\"line-height: 1.714285714; font-size: 1rem;\">The trick is for providers and researchers to keep talking and conquer this \u2018unruly\u2019 Wild West of digital humanities bit by bit, step by step.<\/span><\/p>\n<p>And, by all means, allow researchers to \u2018tinker\u2019 with the data. Verheul: \u2018There is a certain serendipity in working with big data that allows for playfulness.\u2019<\/p>\n<p>See also:<\/p>\n<ul>\n<li><a href=\"https:\/\/gist.github.com\/drjwbaker\/10355412\">James Baker\u2019s (British Library) miniblog<\/a><\/li>\n<li><a href=\"https:\/\/twitter.com\/search?q=%23digrep14&amp;src=tyah\">Twitter hash tag\u00a0#digrep14<\/a><\/li>\n<li>Arjan van Hessen\u2019s blog at\u00a0<a href=\"http:\/\/www.clarin.nl\/\">CLARIN<\/a><\/li>\n<li><a href=\"http:\/\/www.europeana-newspapers.eu\/qa-with-newspapers-researchers-toine-pieters\/\">A\u00a0recent interview with Toine Pieters at the Europeana Newspapers site<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>(Reblogged from Research in KB\u00a0blog:\u00a0http:\/\/researchkb.wordpress.com\/2014\/04\/13\/how-to-maximise-usage-of-digital-collections\/) Libraries want to understand the researchers who use their digital collections and researchers want to understand the nature of these collections better. The seminar \u2018Mining digital repositories\u2019 brought them together at the Dutch Koninklijke Bibliotheek (KB) on 10-11 April, 2014,<\/p>\n","protected":false},"author":4217,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[153],"tags":[178,192],"class_list":["post-2810","post","type-post","status-publish","format-standard","hentry","category-news","tag-digitisation","tag-europeana"],"acf":[],"_links":{"self":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2810","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/users\/4217"}],"replies":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2810"}],"version-history":[{"count":0,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2810\/revisions"}],"wp:attachment":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2810"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2810"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2810"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}