{"id":2824,"date":"2014-06-05T12:17:06","date_gmt":"2014-06-05T12:17:06","guid":{"rendered":"https:\/\/digitisation.eu\/blog\/?p=1910"},"modified":"2014-06-05T12:17:06","modified_gmt":"2014-06-05T12:17:06","slug":"datech-3rd-session-postcorrection","status":"publish","type":"post","link":"https:\/\/digitisation.eu\/?p=2824","title":{"rendered":"DATeCH 3rd Session &#8211; Postcorrection"},"content":{"rendered":"<p>After the lunch break, the DATeCH conference continued with its third session, chaired by Martin Reynaert (Tilburg University), where the Postcorrection of OCR was discussed.<!--more--><\/p>\n<p><strong>John Evershed<\/strong>, from Project Computing\u00a0described a system for automatic post OCR text correction of\u00a0digital collections of historical texts (based on a &#8220;noisy channel&#8221; approach)<br \/>\nwhich avoids manual correction.<\/p>\n<p><iframe loading=\"lazy\" style=\"border: 1px solid #CCC; border-width: 1px 1px 0; margin-bottom: 5px; max-width: 100%;\" src=\"http:\/\/www.slideshare.net\/slideshow\/embed_code\/35422861\" height=\"356\" width=\"427\" allowfullscreen=\"\" frameborder=\"0\" marginwidth=\"0\" marginheight=\"0\" scrolling=\"no\"><\/iframe><\/p>\n<p>Watch the video:<\/p>\n[cvm_video id=&#8221;2093&#8243;]\n<p><strong>G\u00fcnter M\u00fchlberger<\/strong>, from Innsbruck University, introduced a new approach to the correction of\u00a0noisy OCR text which combines the power of crowdsourcing with\u00a0information retrieval technology. It provides a view of the word snippets of a specific\u00a0search string and the possibility of validating each word snippet\u00a0with a simple yes\/no decision.<\/p>\n<p><iframe loading=\"lazy\" style=\"border: 1px solid #CCC; border-width: 1px 1px 0; margin-bottom: 5px; max-width: 100%;\" src=\"http:\/\/www.slideshare.net\/slideshow\/embed_code\/35422860\" height=\"356\" width=\"427\" allowfullscreen=\"\" frameborder=\"0\" marginwidth=\"0\" marginheight=\"0\" scrolling=\"no\"><\/iframe><\/p>\n<p>You can watch the video:<\/p>\n[cvm_video id=&#8221;2072&#8243;]\n<p><strong>Christoph Ringlstetter<\/strong>, from Gini GmbH, presented a new tool which visualizes possible OCR errors and series of similar possible OCR errors in a given input document, and allows therefore for the correction of multiple errors in one shot.<\/p>\n<p><iframe loading=\"lazy\" style=\"border: 1px solid #CCC; border-width: 1px 1px 0; margin-bottom: 5px; max-width: 100%;\" src=\"http:\/\/www.slideshare.net\/slideshow\/embed_code\/35422864\" height=\"356\" width=\"427\" allowfullscreen=\"\" frameborder=\"0\" marginwidth=\"0\" marginheight=\"0\" scrolling=\"no\"><\/iframe><\/p>\n<figure id=\"attachment_1874\" aria-describedby=\"caption-attachment-1874\" style=\"width: 300px\" class=\"wp-caption alignleft\"><a href=\"https:\/\/digitisation.eu\/blog\/wp-content\/uploads\/2014\/06\/Exhibit.jpg\"><img loading=\"lazy\" decoding=\"async\" class=\"size-medium wp-image-1874\" alt=\"The stand of the Spanish company Tecnil\u00f3gica, awarded with a Commendation of Merit of the Succeed Awards, during the Digitisation Days\" src=\"https:\/\/digitisation.eu\/blog\/wp-content\/uploads\/2014\/06\/Exhibit-300x201.jpg\" width=\"300\" height=\"201\" \/><\/a><figcaption id=\"caption-attachment-1874\" class=\"wp-caption-text\">The stand of the Spanish company Tecnil\u00f3gica, awarded with a Commendation of Merit of the Succeed Awards, during the Digitisation Days<\/figcaption><\/figure>\n<p>&nbsp;<\/p>\n<p>At the end of this session, Tecnil\u00f3gica representatives introduced their company, followed by a coffee break at the companies exhibition hall.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>Watch the video:<\/p>\n[cvm_video id=&#8221;2124&#8243;]\n<p>&nbsp;<\/p>\n<p><iframe loading=\"lazy\" style=\"border: 1px solid #CCC; border-width: 1px 1px 0; margin-bottom: 5px; max-width: 100%;\" src=\"http:\/\/www.slideshare.net\/slideshow\/embed_code\/35527790\" height=\"356\" width=\"427\" allowfullscreen=\"\" frameborder=\"0\" marginwidth=\"0\" marginheight=\"0\" scrolling=\"no\"><\/iframe><\/p>\n","protected":false},"excerpt":{"rendered":"<p>After the lunch break, the DATeCH conference continued with its third session, chaired by Martin Reynaert (Tilburg University), where the Postcorrection of OCR was discussed.<\/p>\n","protected":false},"author":4217,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[156,153,176],"tags":[214,218,200,178,186,157,244,185,245],"class_list":["post-2824","post","type-post","status-publish","format-standard","hentry","category-events","category-news","category-succeed","tag-bne","tag-datech","tag-ddays","tag-digitisation","tag-digitisation-days","tag-impact","tag-innsbruck","tag-succeed","tag-tecnilogica"],"acf":[],"_links":{"self":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2824","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/users\/4217"}],"replies":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2824"}],"version-history":[{"count":0,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2824\/revisions"}],"wp:attachment":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2824"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2824"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2824"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}