{"id":3197,"date":"2016-11-10T11:47:26","date_gmt":"2016-11-10T10:47:26","guid":{"rendered":"https:\/\/digitisation.eu\/?p=3197"},"modified":"2017-05-12T13:39:36","modified_gmt":"2017-05-12T11:39:36","slug":"breakthrough-in-archival-access-googling-through-archives-within-reach","status":"publish","type":"post","link":"https:\/\/digitisation.eu\/?p=3197","title":{"rendered":"Breakthrough in archival access: Googling through archives within reach"},"content":{"rendered":"<p style=\"text-align: center;\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-1756\" src=\"https:\/\/digitisation.eu\/wp-content\/uploads\/2015\/02\/logoweb.png\" alt=\"logoweb\" width=\"200\" height=\"90\" \/>\u00a0 \u00a0 \u00a0 \u00a0<img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-3200\" src=\"https:\/\/digitisation.eu\/wp-content\/uploads\/2016\/11\/Ntionaal-archief-logo-300x130.png\" alt=\"Ntionaal-archief-logo\" width=\"200\" height=\"87\" \/>\u00a0<img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-3207\" src=\"https:\/\/digitisation.eu\/wp-content\/uploads\/2016\/11\/NOB-logo-300x179.png\" alt=\"NOB logo\" width=\"200\" height=\"120\" \/>\u00a0<img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-3205 aligncenter\" src=\"https:\/\/digitisation.eu\/wp-content\/uploads\/2016\/11\/CLST-logo-300x75.jpg\" alt=\"CLST logo\" width=\"319\" height=\"80\" \/><\/p>\n<h2>Pilot project Full Automatic Archival Access completed<\/h2>\n<p><b>AMSTERDAM|THE HAGUE|MADRID|NIJMEGEN \u2013 The ability to google through records is within reach, concludes the final report of the project Full Automatic Archival Access (FAAA). This project studied the opportunities to use new digital technologies to make paper based archives searchable at document-level. Four out of five words were correctly recognized by OCR- and NER-software.<\/b><!--more--><\/p>\n<p><span lang=\"en-US\">A small selection of the Central Archive of Justice (CABR, National Archives of the Netherlands) was used in the pilot. Project partners the Network of Dutch War Collections, Centre for Language and Speech Technology, National Archives of the Netherlands and IMPACT Centre of Competence are pleasantly surprised with the results.<\/span><\/p>\n<p><span lang=\"en-US\">Eighty-one percent of the words in the test-documents are correctly recognized by software. That means that it is possible to make typed or hybrid text documents with a standard layout automatically, digitally searchable with an acceptable error rate. A standard layout exists of straight lines, a regular ink density and clear contrast between text and background.<\/span><\/p>\n<p><span lang=\"en-US\">The FAAA-project consisted of two steps. First, the approximately one hundred documents from the CABR-archive have been made machine-readable with use of Optical Character Recognition (OCR)-software. Then, the quality of the OCR\u2019ed-text was improved by using Named Entity Recognition (NER)-software. This software is able to select places, persons and organizations and correct them if necessary.<\/span><\/p>\n<p><span lang=\"en-US\">A leap forward in the accessibility of archives, which are currently mostly described on collection or sub-collection level and rarely accessible on document-level. Program director Network of Dutch War Collections Puck Huitsing: \u201cThe ability to make archives automatically digitally searchable offers many new opportunities for researchers. Historical collections can be questioned in a way that has never been possible in the paper world\u201d. <\/span><\/p>\n<p><span lang=\"en-US\">The project Full Automatic Archival Access was funded by Archief2020, BRAIN, VSBFonds, VFonds and the Ministry of Health, Welfare and Sport. The final and individual reports are published on the Network of Dutch War Collections-website: <\/span><a href=\"http:\/\/oorlogsbronnen.nl\/volauto\"><span style=\"color: #00000a;\"><span lang=\"en-US\">http:\/\/oorlogsbronnen.nl\/volauto<\/span><\/span><\/a><span lang=\"en-US\">.<\/span><\/p>\n<figure id=\"attachment_3198\" aria-describedby=\"caption-attachment-3198\" style=\"width: 300px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-3198 size-medium\" src=\"https:\/\/digitisation.eu\/wp-content\/uploads\/2016\/11\/FAAA-Nov-8-16-Hybrid-document-source-NIOD-300x237.jpg\" alt=\"FAAA Nov 8 16 - Hybrid document (source NIOD)\" width=\"300\" height=\"237\" \/><figcaption id=\"caption-attachment-3198\" class=\"wp-caption-text\">An example of a hybrid document, with printed and handwritten text (source: NIOD Institute for War, Holocaust and Genocide studies).<\/figcaption><\/figure>\n<p><b>Organizations<\/b><\/p>\n<p><span lang=\"en-US\">The National Archives of the Netherlands in The Hague holds 125 kilometres of documents, photos and maps from the central government, and organizations and persons of national importance (past and present).<\/span><\/p>\n<p><span lang=\"en-US\">The Network of Dutch War Collections (NOB) is facilitated by NIOD Institute for War, Holocaust and Genocide Studies. It wants to enhance the use of sources from and about the Second World War in the Netherlands by making the scattered sources better findable and more usable. <\/span><\/p>\n<p><span lang=\"en-US\">The Centre for Language and Speech Technology (CLST) from the Radboud University Nijmegen aims to contribute to the development of language and speech technology. CLST is active in research, application development and consultancy.<\/span><\/p>\n<p><span lang=\"en-US\">The Impact Centre of Competence is a not for profit organisation, comprised of public and private institutions, with the mission to make the digitisation of historical printed text \u2018better, faster and cheaper\u2019. It provides tools, services and facilities to further advance the state-of-the-art in the field of document imaging, language technology and the processing of historical text.<\/span><\/p>\n<p><span lang=\"en-US\"><b>Note for editors<\/b><\/span><\/p>\n<p>If you have any questions or want more information, you can contact Edwin Klijn, program manager Network of Dutch War Collections (edwin.klijn(at)oorlogsbronnen.nl or (+31)020-5233830).<\/p>\n<p style=\"text-align: center;\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-3206 size-medium\" src=\"https:\/\/digitisation.eu\/wp-content\/uploads\/2016\/11\/csm_Logo_VWS_EN_Officieel_NL_8e8efa78f8-300x109.png\" alt=\"csm_Logo_VWS_EN_Officieel_NL_8e8efa78f8\" width=\"300\" height=\"109\" \/> <img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-3209 size-medium\" src=\"https:\/\/digitisation.eu\/wp-content\/uploads\/2016\/11\/vsbfonds-logo-300x66.jpg\" alt=\"VSBfonds PAY-OFF CMYK\" width=\"300\" height=\"66\" \/>\u00a0 \u00a0 \u00a0<img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-3208\" src=\"https:\/\/digitisation.eu\/wp-content\/uploads\/2016\/11\/vfonds-logo-open-engels-300x248.jpg\" alt=\"vfonds-logo-open-engels\" width=\"175\" height=\"145\" \/><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-3204\" src=\"https:\/\/digitisation.eu\/wp-content\/uploads\/2016\/11\/BRAIN-logo.gif\" alt=\"BRAIN logo\" width=\"175\" height=\"142\" \/>\u00a0 \u00a0 \u00a0\u00a0<img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-3203\" src=\"https:\/\/digitisation.eu\/wp-content\/uploads\/2016\/11\/Archief2020-logo-300x206.jpeg\" alt=\"Archief2020 logo\" width=\"212\" height=\"146\" \/><\/p>\n","protected":false},"excerpt":{"rendered":"<p>\u00a0 \u00a0 \u00a0 \u00a0\u00a0\u00a0 Pilot project Full Automatic Archival Access completed AMSTERDAM|THE HAGUE|MADRID|NIJMEGEN \u2013 The ability to google through records is within reach, concludes the final report of the project Full Automatic Archival Access (FAAA). This project studied the opportunities to use new digital technologies to make paper based archives searchable at document-level. Four out &hellip; <a href=\"https:\/\/digitisation.eu\/?p=3197\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Breakthrough in archival access: Googling through archives within reach&#8221;<\/span><\/a><\/p>\n","protected":false},"author":476,"featured_media":12041,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[153],"tags":[],"class_list":["post-3197","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-news"],"acf":[],"_links":{"self":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/3197","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/users\/476"}],"replies":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=3197"}],"version-history":[{"count":1,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/3197\/revisions"}],"predecessor-version":[{"id":12039,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/3197\/revisions\/12039"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/media\/12041"}],"wp:attachment":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=3197"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=3197"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=3197"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}