{"id":2763,"date":"2011-10-25T11:33:11","date_gmt":"2011-10-25T11:33:11","guid":{"rendered":"http:\/\/impactocr.wordpress.com\/?p=989"},"modified":"2011-10-25T11:33:11","modified_gmt":"2011-10-25T11:33:11","slug":"impact-final-conference-overview-of-language-work-in-impact","status":"publish","type":"post","link":"https:\/\/digitisation.eu\/?p=2763","title":{"rendered":"IMPACT Final Conference-Overview of language work in IMPACT"},"content":{"rendered":"<figure id=\"attachment_1032\" aria-describedby=\"caption-attachment-1032\" style=\"width: 300px\" class=\"wp-caption alignleft\"><a href=\"https:\/\/digitisation.eu\/blog\/wp-content\/uploads\/2011\/10\/impact-conf-d-49.jpg\"><img loading=\"lazy\" decoding=\"async\" class=\"size-medium wp-image-1032\" title=\"Katrien  Depuydt\" alt=\"Katrien  Depuydt\" src=\"https:\/\/digitisation.eu\/blog\/wp-content\/uploads\/2011\/10\/impact-conf-d-49.jpg?w=300\" width=\"300\" height=\"231\" \/><\/a><figcaption id=\"caption-attachment-1032\" class=\"wp-caption-text\">Katrien Depuydt gives an overview of language work in IMPACT<\/figcaption><\/figure>\n<p>Katrien Depuydt provided a brief overview of the IMPACT project\u2019s work packages devoted to creating language tools and lexicon to aid in both information retrieval and OCR processing. How might one measure successful improvement to the access of text? She cleverly posits that the key will be in asking ourselves: \u00a0\u201cCan we handle the \u201cworld\u201d?<\/p>\n<p>&nbsp;<\/p>\n<p><!--more--><\/p>\n<p>In an 18<sup>th <\/sup>century dutch periodical \u2018werried\u2019 was the spelling of the day and using OCR built with a simple dutch dictionary you would need to begin your search with that term and the results would be necessarily limited. What we really want she notes, is to key in the modern term\u00a0\u201cworld\u201d and retrieve all the appropriate variants in the text.<\/p>\n<figure id=\"attachment_1033\" aria-describedby=\"caption-attachment-1033\" style=\"width: 300px\" class=\"wp-caption alignright\"><a href=\"https:\/\/digitisation.eu\/blog\/wp-content\/uploads\/2011\/10\/impact-conf-d-41.jpg\"><img loading=\"lazy\" decoding=\"async\" class=\"size-medium wp-image-1033\" title=\"Katrien Dupuydt\" alt=\"Katrien Dupuydt\" src=\"https:\/\/digitisation.eu\/blog\/wp-content\/uploads\/2011\/10\/impact-conf-d-41.jpg?w=300\" width=\"300\" height=\"214\" \/><\/a><figcaption id=\"caption-attachment-1033\" class=\"wp-caption-text\">Katrien Dupuydt gives an overview of language work in in IMPACT<\/figcaption><\/figure>\n<p>This is where IMPACT\u2019s work in building lexica comes in, and we start to\u00a0discover that yes, we\u00a0CAN handle \u2018the world\u2019. In the course of the project an <a href=\"http:\/\/impactocr.wordpress.com\/2011\/07\/12\/lexical-tools-with-katrien-depuydt\/\" target=\"_blank\">OCR lexicon, an IR lexicon and an NE lexicon <\/a>were created for 9 languages and these plug into ABBYY FineReader enhancing the OCR and the retrieval. No simple task, the work required analysing different language resources available for each unique language, identifying tools already available, special character sets and the like. She gives the example of Bulgarian which had no existing dictionaries or lexica and some characters were not recognised by Abbey Fine Reader creating a unique set of challenges. How these challenges were overcome will be explored in more detail at the forthcoming in the Parallel Session 2: Language Session later this afternoon.<\/p>\n<p>View the presentation here:<br \/>\n[slideshare id=9875684&amp;doc=katriendepuydt-111025105716-phpapp02]\n<p>and the video here:<br \/>\nhttp:\/\/www.vimeo.com\/32505214<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Katrien Depuydt provided a brief overview of the IMPACT project\u2019s work packages devoted to creating language tools and lexicon to aid in both information retrieval and OCR processing. How might one measure successful improvement to the access of text? She cleverly posits that the key will be in asking ourselves: \u00a0\u201cCan we handle the \u201cworld\u201d? &hellip; <a href=\"https:\/\/digitisation.eu\/?p=2763\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;IMPACT Final Conference-Overview of language work in IMPACT&#8221;<\/span><\/a><\/p>\n","protected":false},"author":4217,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[154,161],"tags":[],"class_list":["post-2763","post","type-post","status-publish","format-standard","hentry","category-discussions","category-ocr-dictionaries-and-lexica"],"acf":[],"_links":{"self":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2763","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/users\/4217"}],"replies":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2763"}],"version-history":[{"count":0,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2763\/revisions"}],"wp:attachment":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2763"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2763"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2763"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}