{"id":2693,"date":"2010-05-07T12:19:12","date_gmt":"2010-05-07T12:19:12","guid":{"rendered":"http:\/\/impactocr.wordpress.com\/?p=174"},"modified":"2010-05-07T12:19:12","modified_gmt":"2010-05-07T12:19:12","slug":"an-overview-of-technical-solutions-in-impact","status":"publish","type":"post","link":"https:\/\/digitisation.eu\/?p=2693","title":{"rendered":"The challenges of historical materials and an overview on the technical solutions in IMPACT"},"content":{"rendered":"<p>Sven Schlarb of the \u00d6sterreichische Nationalbibliothek (Austrian National Library) now talks about the challenges of text digitisation for OCR and the solutions IMPACT has devised to deal with them.\u00a0 Having outlined the individual tools and the partners responsible for their development, he talks in detail about the ideal IMPACT workflow in which they can all be used.<\/p>\n<p><!--more--><\/p>\n[slideshare id=4138311&amp;doc=bratislavaws-schlarb-onb-technicaltools-100518090407-phpapp01]\n[vimeo http:\/\/vimeo.com\/11650707]\n<p>One important novelty to the IMPACT approach is that it begins with image enhancement (border detection, geometric correction and binarisation) before the page is segmented into text blocks.\u00a0 This links to a point that both Aly and G\u00fcnter have made this morning: that a huge amount of digitised material exists that was not captured with OCR in mind.\u00a0 Enhancing images upfront is done to translate those images into something like an optimal scan for OCR.<\/p>\n<p>Having discussed the various tools in depth, Sven talks about interoperability and sharing of results from one tool to another.\u00a0 IMPACT is using xml based standards for all OCR outputs, combining them in an &#8220;embracing&#8221; or translation standard called PAGE xml, being developed by the University of Salford.<\/p>\n<p><em>Niall Anderson, BL + Mark-Oliver Fischer, BSB<br \/>\n<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Sven Schlarb of the \u00d6sterreichische Nationalbibliothek (Austrian National Library) now talks about the challenges of text digitisation for OCR and the solutions IMPACT has devised to deal with them.\u00a0 Having outlined the individual tools and the partners responsible for their development, he talks in detail about the ideal IMPACT workflow in which they can all &hellip; <a href=\"https:\/\/digitisation.eu\/?p=2693\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;The challenges of historical materials and an overview on the technical solutions in IMPACT&#8221;<\/span><\/a><\/p>\n","protected":false},"author":4217,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[154,159],"tags":[],"class_list":["post-2693","post","type-post","status-publish","format-standard","hentry","category-discussions","category-optical-character-recognition"],"acf":[],"_links":{"self":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2693","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/users\/4217"}],"replies":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2693"}],"version-history":[{"count":0,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2693\/revisions"}],"wp:attachment":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2693"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2693"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2693"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}