{"id":2809,"date":"2014-04-22T09:08:17","date_gmt":"2014-04-22T09:08:17","guid":{"rendered":"https:\/\/digitisation.eu\/blog\/?p=1743"},"modified":"2014-04-22T09:08:17","modified_gmt":"2014-04-22T09:08:17","slug":"succeed-2nd-hackathon","status":"publish","type":"post","link":"https:\/\/digitisation.eu\/?p=2809","title":{"rendered":"Working together to improve text digitisation techniques"},"content":{"rendered":"<p><strong>2nd Succeed hackathon at the University of Alicante<\/strong><\/p>\n<p>Is there anyone out there\u00a0<span style=\"font-size: 1rem;\">still<\/span><span style=\"font-size: 1rem;\">\u00a0<\/span><span style=\"line-height: 1.714285714; font-size: 1rem;\">thinking that a hackathon is a malicious break-in? <\/span><\/p>\n<p><span style=\"line-height: 1.714285714; font-size: 1rem;\">Far from it.\u00a0<\/span><span lang=\"NL\" style=\"line-height: 1.714285714; font-size: 1rem;\">It is the best way for developers and researchers to get together and work on new tools and innovations. The<\/span><a style=\"line-height: 1.714285714; font-size: 1rem;\" href=\"http:\/\/impact.dlsi.ua.es\/wiki\/index.php\/Developers_workshops_(hackathons)#2014\" target=\"_blank\">2nd developers workshop \/ hackathon<\/a><span style=\"line-height: 1.714285714; font-size: 1rem;\">\u00a0organised on 10-11 April by the <\/span><a style=\"line-height: 1.714285714; font-size: 1rem;\" href=\"http:\/\/succeed-project.eu\/\" target=\"_blank\">Succeed Project<\/a>\u00a0<span style=\"line-height: 1.714285714; font-size: 1rem;\">was a case in point: bringing together people to work on new ideas and new inspiration for better OCR.\u00a0The event was held in the Claude Shannon room of the Department of Software\u00a0and Computing Systems (<\/span><a style=\"line-height: 1.714285714; font-size: 1rem;\" href=\"http:\/\/www.dlsi.ua.es\/index.cgi?id=eng\" target=\"_blank\">DLSI<\/a><span style=\"line-height: 1.714285714; font-size: 1rem;\">) of the <\/span><a style=\"line-height: 1.714285714; font-size: 1rem;\" href=\"http:\/\/web.ua.es\/en\/actualidad-universitaria\/2014\/abril2014\/abril2014-7-13\/la-ua-acoge-el-segundo-hackathon-del-proyecto-succeed.html\" target=\"_blank\">University of Alicante<\/a><span style=\"line-height: 1.714285714; font-size: 1rem;\">, Spain.\u00a0<\/span><a style=\"line-height: 1.714285714; font-size: 1rem;\" href=\"http:\/\/en.wikipedia.org\/wiki\/Claude_Shannon\" target=\"_blank\">Claude Shannon<\/a><span style=\"line-height: 1.714285714; font-size: 1rem;\">\u00a0was a famous mathematician and engineer and is also known as the &#8220;father of information theory&#8221;. So it seemed a good place to have a hackathon!<\/span><\/p>\n<p><!--more--><\/p>\n<p><a href=\"https:\/\/www.flickr.com\/photos\/116354723@N02\/13757270124\/\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/live.staticflickr.com\/7176\/13757270124_18bd244d9d_c.jpg\" alt=\"Ready to start 2\" width=\"800\" height=\"468\" \/><\/a><\/p>\n<p><iframe loading=\"lazy\" title=\"Succeed Second Developer&#039;s  Workshop - Hackathon\" width=\"840\" height=\"473\" src=\"https:\/\/www.youtube.com\/embed\/oQ4k3TLdc9s?feature=oembed\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share\" allowfullscreen><\/iframe><script type=\"application\/json\" data-arve-oembed>{\"title\":\"Succeed Second Developer&#039;s  Workshop - Hackathon\",\"author_name\":\"IMPACT Centre of Competence\",\"author_url\":\"https:\/\/www.youtube.com\/@ImpactCentreofCompetence\",\"type\":\"video\",\"height\":\"473\",\"width\":\"840\",\"version\":\"1.0\",\"provider_name\":\"YouTube\",\"provider_url\":\"https:\/\/www.youtube.com\/\",\"thumbnail_height\":\"360\",\"thumbnail_width\":\"480\",\"thumbnail_url\":\"https:\/\/i.ytimg.com\/vi\/oQ4k3TLdc9s\/hqdefault.jpg\",\"html\":\"&lt;iframe width=&quot;840&quot; height=&quot;473&quot; src=&quot;https:\/\/www.youtube.com\/embed\/oQ4k3TLdc9s?feature=oembed&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen title=&quot;Succeed Second Developer&#039;s  Workshop - Hackathon&quot;&gt;&lt;\/iframe&gt;\",\"arve_cachetime\":\"2024-03-03 23:44:18\",\"arve_url\":\"https:\/\/www.youtube.com\/watch?v=oQ4k3TLdc9s\",\"arve_srcset\":\"https:\/\/i.ytimg.com\/vi\/oQ4k3TLdc9s\/mqdefault.jpg 320w, https:\/\/i.ytimg.com\/vi\/oQ4k3TLdc9s\/hqdefault.jpg 480w, https:\/\/i.ytimg.com\/vi\/oQ4k3TLdc9s\/sddefault.jpg 640w, https:\/\/i.ytimg.com\/vi\/oQ4k3TLdc9s\/maxresdefault.jpg 1280w\"}<\/script><\/p>\n<p><em>Clemens explains what a hackathon is and what we hope to achieve with it for Succeed.<\/em><\/p>\n<p>Same as <a href=\"https:\/\/digitisation.eu\/blog\/1st-succeed-hackathon-kb\/\" target=\"_blank\">last year<\/a>, we again provided a <a href=\"http:\/\/impact.dlsi.ua.es\/wiki\/index.php\/Developers_workshops_(hackathons)#2014\" target=\"_blank\">wiki<\/a>\u00a0upfront with some information about possible topics to work on, as well as a number of tools and data that participants could experiment with before and\u00a0during the event. Unfortunately there was an unexpectedly high number of no-shows this time &#8211; we try to keep these events free and open to everyone, but may have to think about charging at least a no-show fee in the future, as places are usually limited. Did those hackers simply have to stay home to fix the\u00a0<a href=\"http:\/\/xkcd.com\/1354\/\" target=\"_blank\">heartbleed<\/a>\u00a0bug on their servers? We will probably never find out.<\/p>\n<h3>Collaboration, open source tools, open solutions<\/h3>\n<p>Nevertheless, there was a large enough group of programmers and researchers from Germany, Poland, the Netherlands, and various parts of Spain eager to immerse themselves\u00a0deeply into a diverse list of\u00a0topics. Already in the introduction we agreed to work on open tools and solutions, and quickly identified some areas in which open source tool support for text digitisation is still lacking (see below). Actually, one of the first things we did was to set up a local <a href=\"http:\/\/git-scm.com\/\" target=\"_blank\">git<\/a> repository, and people were pushing code samples, prototypes, and interesting projects to share with the group during both\u00a0days.<\/p>\n<p><a href=\"https:\/\/www.flickr.com\/photos\/116354723@N02\/13775777003\/\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/live.staticflickr.com\/7417\/13775777003_d85cef2fb1_c.jpg\" alt=\"Second Day, April 11_5\" width=\"800\" height=\"511\" \/><\/a><\/p>\n<h3>What&#8217;s the status of open source OCR?<\/h3>\n<p>In answer to that question, Jes\u00fas Dominguez Muriel from\u00a0<a href=\"http:\/\/www.digibis.com\/\" target=\"_blank\">Digib\u00eds<\/a>\u00a0(the company that also made\u00a0<a href=\"http:\/\/www.digibis.com\/dpla-europeana\/\" target=\"_blank\">http:\/\/www.digibis.com\/dpla-europeana\/<\/a>) started\u00a0an\u00a0investigation into open source OCR tools and frameworks. He made a really detailed\u00a0analysis of the status of open source OCR, which you can find\u00a0<a href=\"https:\/\/github.com\/impactcentre\/hackathon2014\/raw\/master\/slides\/jdmuriel.pdf\" target=\"_blank\">here<\/a>. Thanks a lot for that summary, Jes\u00fas! At\u00a0the end of his presentation, Jes\u00fas also suggested an &#8220;algorithm wikipedia&#8221; &#8211; I guess something similar to <a href=\"http:\/\/rosettacode.org\/wiki\/Rosetta_Code\" target=\"_blank\">RosettaCode<\/a> but then specifically for OCR. This would indeed be very useful to share algorithms but also implementations and prevent reinventing (or reimplementing) the wheel. Something for our new\u00a0<a href=\"http:\/\/impact.dlsi.ua.es\/wiki\/index.php\/OCRpedia\" target=\"_blank\">OCRpedia<\/a>, perhaps?<\/p>\n<h3>A method for assessing OCR quality based on ngrams<\/h3>\n<p>As it turned out on the second day, a very promising idea seemed to be using <a href=\"http:\/\/en.wikipedia.org\/wiki\/N-gram\" target=\"_blank\">ngrams<\/a> for assessing the quality of an OCR&#8217;ed text, without the need for ground truth. Well, in fact you do still need some correct text to create the ngram model, but one can use texts from e.g. <a href=\"http:\/\/www.gutenberg.org\/\" target=\"_blank\">Project Gutenberg<\/a>\u00a0or <a href=\"http:\/\/aspell.net\/\" target=\"_blank\">aspell<\/a> for that. Two groups started to work on this:\u00a0while <a href=\"https:\/\/twitter.com\/willemjanfaber\" target=\"_blank\">Willem Jan Faber<\/a> from the KB experimented with\u00a0a simple <a href=\"https:\/\/github.com\/impactcentre\/hackathon2014\/blob\/master\/ngram-ocr-eval\/ngram_ocr.py\" target=\"_blank\">Python script<\/a>\u00a0for that purpose, the group of Rafael Carrasco, Sebastian Kirch and Tomasz Parkola decided to implement this as a new feature in the Java <a href=\"https:\/\/github.com\/impactcentre\/ocrevalUAtion\" target=\"_blank\">ocrevalUAtion<\/a> tool (check the work-in-progress &#8220;<a href=\"https:\/\/github.com\/impactcentre\/ocrevalUAtion\/tree\/wip\" target=\"_blank\">wip<\/a>&#8221; branch).<\/p>\n<p><a href=\"https:\/\/www.flickr.com\/photos\/116354723@N02\/13775774723\/\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/live.staticflickr.com\/3757\/13775774723_e2599371fd_c.jpg\" alt=\"Second Day, April 11_4\" width=\"800\" height=\"600\" \/><\/a><\/p>\n<p><em>Jes\u00fas in the front, Rafael, Sebastian and Tomasz discussing ngrams in the back.<\/em><\/p>\n<h3>Aligning text and segmentation results<\/h3>\n<p>Another very promising development was started by Antonio Corbi from the University of Alicante. He worked on a software\u00a0to align plain text and segmentation results. The idea is to first identify all the lines in a document, segment them into words and eventually individual charcaters, and then align the character outlines with the text in the ground truth. This would allow (among other things) creating a large corpus of training material for an OCR classifier based on\u00a0the more than <a href=\"http:\/\/www.primaresearch.org\/datasets\/IMPACT_Digitisation\" target=\"_blank\">50,000 images with ground truth<\/a> produced in the <a href=\"http:\/\/www.impact-project.eu\/\" target=\"_blank\">IMPACT Project<\/a>, for which correct text is available, but segmentation could only be done on\u00a0the level of regions. Another great feature of <a href=\"https:\/\/github.com\/impactcentre\/GroundTruthAligner\" target=\"_blank\">Antonio&#8217;s tool<\/a> is that while he uses <a href=\"http:\/\/dlang.org\/\" target=\"_blank\">D<\/a> as a programming language, he also makes use of <a href=\"http:\/\/www.gtk.org\/\" target=\"_blank\">GTK<\/a>, which has the nice effect that his tool does not only work on the\u00a0desktop, but also as a web application in a browser.<\/p>\n<p><a href=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/aligner.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-large wp-image-934\" alt=\"aligner\" src=\"http:\/\/researchkb.files.wordpress.com\/2014\/04\/aligner.png?w=640\" width=\"640\" height=\"457\" \/><\/a><\/p>\n<h3>OCR is complicated, but don&#8217;t worry\u00a0&#8211;\u00a0we&#8217;re on it!<\/h3>\n<p>Gustavo Candela works for the <a href=\"http:\/\/www.cervantesvirtual.com\/\" target=\"_blank\">Biblioteca Virtual\u00a0Miguel de Cervantes<\/a>, the largest Digital Library in the Spanish speaking world. Usually he is busy with Linked Data and things like\u00a0<a href=\"http:\/\/en.wikipedia.org\/wiki\/Functional_Requirements_for_Bibliographic_Records\" target=\"_blank\">FRBR<\/a>, so he was happy to expand his knowledge and learn about the various processes involved in OCR and what tools and standards are commonly used. His findings: there\u00a0is a lot more complexity involved in OCR than appears at first sight. And again, for some problems it would be good to have more open source tool support.<\/p>\n<p>In fact, at the same time as the hackathon,\u00a0at the KB in The Hague, the &#8216;<a href=\"http:\/\/researchkb.wordpress.com\/2014\/04\/13\/how-to-maximise-usage-of-digital-collections\/\" target=\"_blank\">Mining Digital Repositories<\/a>&#8216; conference was going on where\u00a0the problem of bad OCR was discussed from a scholarly perspective. And also there, the need for more open technologies and methods was apparent:<\/p>\n<blockquote class=\"twitter-tweet\" data-width=\"550\" data-dnt=\"true\">\n<p lang=\"en\" dir=\"ltr\"><a href=\"https:\/\/twitter.com\/hashtag\/digrep14?src=hash&amp;ref_src=twsrc%5Etfw\">#digrep14<\/a> I&#39;m strangely unconcerned about poor OCR. So long as scholars know the quality, they should know what research can &amp; can&#39;t be done<\/p>\n<p>&mdash; James Baker Battles the Pink SPARQL and QS Robots (@j_w_baker) <a href=\"https:\/\/twitter.com\/j_w_baker\/status\/454523523634315264?ref_src=twsrc%5Etfw\">April 11, 2014<\/a><\/p><\/blockquote>\n<p><script async src=\"https:\/\/platform.twitter.com\/widgets.js\" charset=\"utf-8\"><\/script><\/p>\n<blockquote class=\"twitter-tweet\" lang=\"en\" data-conversation=\"none\"><p><a href=\"https:\/\/twitter.com\/j_w_baker\">@j_w_baker<\/a> Fully ACK. But they need open tools and transparent methods for that. We&#8217;re on it! <a href=\"https:\/\/t.co\/CI3m64fnb0\">https:\/\/t.co\/CI3m64fnb0<\/a><\/p>\n<p>\u2014 Clemens Neudecker (@cneudecker) <a href=\"https:\/\/twitter.com\/cneudecker\/statuses\/454528200572682241\">April 11, 2014<\/a><\/p><\/blockquote>\n<p>&nbsp;<\/p>\n<h3>Open source border detection<\/h3>\n<p>One of the many technologies for text digitisation that are available in the <a href=\"https:\/\/digitisation.eu\/\" target=\"_blank\">IMPACT Centre of Competence<\/a> for image pre-processing is <a href=\"https:\/\/digitisation.eu\/tools\/browse\/image-enhancement\/ncsr-border-detection-and-removal\/\" target=\"_blank\">Border Removal<\/a>. This technique is typically applied\u00a0to remove black borders in a digital image that has been captured while scanning a document. The borders don&#8217;t contain any information, yet they take up expensive storage space, so removing the borders without removing any other relevant information from a scanned document page is a desirable thing to do. However, there is no simple open source tool or implementation for doing that at the moment. So Daniel Torregrosa from the University of Alicante started to research the topic. After some quick experiments with tools like\u00a0<a href=\"http:\/\/www.imagemagick.org\/\" target=\"_blank\">imagemagick<\/a> and <a href=\"https:\/\/www.flameeyes.eu\/projects\/unpaper\" target=\"_blank\">unpaper<\/a>, he eventually decided to work on his own algorithm. You can find the source\u00a0<a href=\"https:\/\/github.com\/impactcentre\/hackathon2014\/blob\/master\/paragraph-detection\/smudge.py\" target=\"_blank\">here<\/a>. Besides, he probably earns the award for the best slide in a\u00a0presentation&#8230;showing us two black pixels on a white background!<\/p>\n<h3>A great venue<\/h3>\n<p>All in all, I think we can really be quite happy with these results.\u00a0And indeed the University of Alicante also did\u00a0a great job hosting us &#8211; there was an excellent internet connection available via cable and wifi, plenty of space and tables to discuss in groups and we were distant enough from the classrooms not to be disturbed by the students or vice versa. Also at any time\u00a0there was excellent and light Spanish food &#8211; Gazpacho, Couscous with vegetables, assorted Montaditos, fresh fruit&#8230; you don&#8217;t make hackers happy with just pizza anymore! Of course there were also ice-cooled drinks and hot\u00a0coffee, and rumours spread that there were also some (alcohol-free?) beers in the cooler, but (un)fortunately there is no evidence of that&#8230;<\/p>\n<h3>To be continued!<\/h3>\n<p>If you want to try out any of the software yourself, just visit\u00a0our\u00a0<a href=\"https:\/\/github.com\/impactcentre\/hackathon2014\" target=\"_blank\">github<\/a>\u00a0and have go! Make sure to also take\u00a0a look at the videos that were made with participants\u00a0<a href=\"https:\/\/www.youtube.com\/watch?v=S2Mlwfr6z9k\" target=\"_blank\">Jes\u00fas<\/a>, <a href=\"https:\/\/www.youtube.com\/watch?v=3xdn5coETek\" target=\"_blank\">Sebastian<\/a>\u00a0and\u00a0<a href=\"https:\/\/www.youtube.com\/watch?v=eH8p4k2VUeo\" target=\"_blank\">Tomasz<\/a>, explaining their intentions and expectations for\u00a0the hackathon. And at the next hackathon, maybe we can welcome you too amongst the participants?<\/p>\n","protected":false},"excerpt":{"rendered":"<p>2nd Succeed hackathon at the University of Alicante Is there anyone out there\u00a0still\u00a0thinking that a hackathon is a malicious break-in? Far from it.\u00a0It is the best way for developers and researchers to get together and work on new tools and innovations. The2nd developers workshop \/ hackathon\u00a0organised on 10-11 April by the Succeed Project\u00a0was a case &hellip; <a href=\"https:\/\/digitisation.eu\/?p=2809\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Working together to improve text digitisation techniques&#8221;<\/span><\/a><\/p>\n","protected":false},"author":4217,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[156,176],"tags":[164,191,160,185,158],"class_list":["post-2809","post","type-post","status-publish","format-standard","hentry","category-events","category-succeed","tag-evaluation","tag-hackathon","tag-ocr","tag-succeed","tag-workshop"],"acf":[],"_links":{"self":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2809","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/users\/4217"}],"replies":[{"embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2809"}],"version-history":[{"count":0,"href":"https:\/\/digitisation.eu\/index.php?rest_route=\/wp\/v2\/posts\/2809\/revisions"}],"wp:attachment":[{"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2809"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2809"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/digitisation.eu\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2809"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}