{"id":2698,"date":"2010-05-07T10:17:13","date_gmt":"2010-05-07T10:17:13","guid":{"rendered":"http:\/\/impactocr.wordpress.com\/?p=164"},"modified":"2010-05-07T10:17:13","modified_gmt":"2010-05-07T10:17:13","slug":"ocr-in-libraries-some-practical-remarks","status":"publish","type":"post","link":"https:\/\/digitisation.eu\/?p=2698","title":{"rendered":"OCR in libraries \u2013 some practical\u00a0remarks"},"content":{"rendered":"<p>Next, G\u00fcnter M\u00fchlberger of Universit\u00e4ts- und Landesbibliothek Tirol (University and Regional Library Tyrol) talks about the situation with regard to full-text digitisation in Europe and elsewhere.\u00a0 He notes that in Europe, OCR has not always been a standard part of the workflow, so there is much legacy digital material that has never been OCR&#8217;d &#8211; and more importantly not created for OCR.\u00a0 He notes, however, that Americans have been more proactive in this field, citing JSTOR and Google Books as examples.<\/p>\n<p><!--more--><\/p>\n[slideshare id=4138337&amp;doc=bratislavaws-mhlberger-ocrinlibrariesnolocallinks-100518090530-phpapp01]\n[vimeo http:\/\/vimeo.com\/11650671]\n<p>There are a few reasons for this, among them the relative unsophistication of even fairly recent OCR systems, but also that a project that involves OCR breeds complications.\u00a0 Put simply: once you&#8217;ve made your OCR, what do you do with it?\u00a0 How do you assure its quality?\u00a0 How do you expose it to the public, if at all?\u00a0 G\u00fcnter gives an example from Austrian Literature Online.<\/p>\n<p>He then outlines the pros and cons of the three major sources of text-based digital material: bound volumes, microfilms and loose folios.\u00a0 All can be done well, but digitisation managers need to focus on producing not just a &#8220;good&#8221; image, but an image that&#8217;s good for OCR.<\/p>\n<p>The characteristics of a good image for OCR are overall sharpness, distinct fonts, a clear background, a complete shot with a white frame around the side for OCR orientation.\u00a0 All lines should be parallel to each other and to page margins.\u00a0 No additional noise (or marginalia) from users.\u00a0 Not all of these things are easy to control, but the existence of problematic features could be used to guide material to be digitised.<\/p>\n<p>Some general recommendations: if it&#8217;s a modern, clean document, it can be captured at 300dpi as a bitonal JPEG; if it&#8217;s an older document with problematic features, then greyscale 400ppi TIFF is preferable. G\u00fcnter concludes by saying that OCR is a must for a digital library text collection.<\/p>\n<p><em>Niall Anderson, The British Library + Mark-Oliver Fischer, Bavarian State Library<br \/>\n<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Next, G\u00fcnter M\u00fchlberger of Universit\u00e4ts- und Landesbibliothek Tirol (University and Regional Library Tyrol) talks about the situation with regard to full-text digitisation in Europe and elsewhere.\u00a0 He notes that in Europe, OCR has not always been a standard part of the workflow, so there is much legacy digital material that has never been OCR&#8217;d &#8211; &hellip; <a href=\"https:\/\/digitisation.eu\/?p=2698\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;OCR in libraries \u2013 some practical\u00a0remarks&#8221;<\/span><\/a><\/p>\n","protected":false},"author":4217,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[154,159],"tags":[160],"class_list":["post-2698","post","type-post","status-publish","format-standard","hentry","category-discussions","category-optical-character-recognition","tag-ocr"],"acf":[],"_links":{"self":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2698","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/users\/4217"}],"replies":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2698"}],"version-history":[{"count":0,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2698\/revisions"}],"wp:attachment":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2698"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2698"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2698"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}