Category: Optical Character Recognition
The IMPACT Framework – From Tools to Workflows
This practical session started with the attendees introducing themselves and splitting up into 3 groups, so that each could work on a different set of tasks based on a Case Study.
IMPACT Final Conference – IBM Adaptive OCR Engine and CONCERT Cooperative Correction
Asaf Tzadok (IBM Haifa Research Lab) showed us IBM’s CONCERT tool which facilitates collaborative OCR correction. CONCERT (Cooperative Engine for the Correction of Extracted Text) works in three steps: character session, word session and page-level session.
Veranstaltungsende
The first part of the OCR workshop in Munich concluded with perspectives of the project IMPACT. The second and more practical orientated part of the OCR workshop can be tracked via the blog of the Munich Centre of Digitisation. Many thanks to the lecturers and sponsors who enabled the exchange of information.
Kollaborative Korrektur
Doris Škarić from the Bavarian State Library reported about collaborative correction of OCR results by volunteers. She presented the IMPACT tool CONCERT (the COllaborative eNgine for the CorREction of Texts) and reported about the findings of a pilot test of the tool at the Bavarian State Library.
Dokumentstrukturerkennung
Günter Mühlberger from the University and Regional Library of Tyrol in Innsbruck presented the Functional Extension Parser (FEP), a tool for the OCR-based structural analysis of printed texts.
Verbesserte OCR-Software für historische Dokumente
In his second talk of the day, Gerd Zechmeister of the Austrian National Library spoke about Optical Character Recognition (OCR), the processing steps of a typical OCR software, and Abbyy’s role as technology provider in IMPACT.
Post processing and language technology in OCR
Jesse De Does from INL follows Katrien by introducing a post-processing Text and Error Profiling tool that looks similar to the IBM tool demonstrated by Niall earlier. This tool differs through the use of ‘text profiling’, whereby language and text is analysed on the basis of frequency and logic. It also allows for batch correction … Continue reading "Post processing and language technology in OCR"
IMPACT Online Resources
Neil Fitzgerald gave a clear overview of the materials that will be provided by the IMPACT project through the new website, including the: Decision Support Tools Learning Resource Toolkit Online Learning Resources (OER available as Xerte Resources on CC licenses) Help Desk
The OCR process – Clemens Neudecker & Niall Anderson
IMPACT work with ABBYY IMPACT has been working with both ABBYY and IBM on OCR and Clemens starts with outlining all the steps required during the OCRing process and the extension of the FineReader Engine so it performs better with historical text.
Image Enhancement talk from Niall Anderson
Niall took us through an overview of why Image Enhancement was necessary for OCR use and what the current state of the art is as practiced by the IMPACT partners.
