BSB Demo Day: Impressionen und überarbeitete Artikel


One week after our very successful Demo Day at the Bavarian State Library, we had a look at all articles and reworded them were necessary, to weed out factual, grammatical and spelling errors. Turns out blogging live just isn’t that easy. Also, the presentation slides were added to all talks. As promised, we also caught … Continue reading "BSB Demo Day: Impressionen und überarbeitete Artikel"

Veranstaltungsende


The first part of the OCR workshop in Munich concluded with perspectives of the project IMPACT. The second and more practical orientated part of the OCR workshop can be tracked via the blog of the Munich Centre of Digitisation. Many thanks to the lecturers and sponsors who enabled the exchange of information.

Centre of Competence


Hildelies Balk – Pennington de Jongh from the Koninklijke Bibliotheek, the National Library of the Netherlands, and project leader of the IMPACT project, outlined the Centre of Competence and the services it will offer. The CoC will be tasked with sustaining the developments of the IMPACT project into the future, and will officially be opened at the … Continue reading "Centre of Competence"

Evaluierungswerkzeuge


Stefan Pletschacher from the University of Salford presented methods to evaluate OCR results. For a proper evaluation, ‘ground truth’, that is almost 100% correct text is needed. But a big challenge lies in how to calculate the gravity of different kinds of errors. Will character or word accuracy be used? Do errors in heading count … Continue reading "Evaluierungswerkzeuge"

Kollaborative Korrektur


Doris Škarić from the Bavarian State Library reported about collaborative correction of OCR results by volunteers. She presented the IMPACT tool CONCERT (the COllaborative eNgine for the CorREction of Texts) and reported about the findings of  a pilot test of the tool at the Bavarian State Library.

Dokumentstrukturerkennung


Günter Mühlberger from the University and Regional Library of Tyrol in Innsbruck presented the Functional Extension Parser (FEP), a tool for the OCR-based structural analysis of printed texts.

Analyse und Nachkorrektur von OCR-Ergebnissen


Ulrich Reffle, who works at the Centre of Information and Language Processing of the Ludwig-Maximilians-University Munich, spoke about document-centric analysis and error detection, which can enable faster and easier correction of OCRed historical texts.

Verbesserte OCR-Software für historische Dokumente


In his second talk of the day, Gerd Zechmeister of the Austrian National Library spoke about Optical Character Recognition (OCR), the processing steps of a typical OCR software, and Abbyy’s role as technology provider in IMPACT.

Spezial-Lexika zur Erschließung historischer Dokumente


Annette Gotscharek of the Centre for Information and Language  Processing of the Ludwig-Maximilians-University talked about special dictionaries and their impact on OCR. Spelling variations and no longer used words are a big problem for OCR, making a lexicon adapted to historical texts a necessity for the creation of correct full text.

Werkzeuge zur Bildoptimierung


Gerd Zechmeister of the Austrian National Library talked about different approaches to improve OCR accuracy by working on the many challenges inherent in the digital images of historical texts (e.g. geometrical correction against curves and creases; removal of black borders around pages, better binarisation, …).