Tag: OCR
IMPACT Final Conference – Evaluation of lexicon supported OCR and information retrieval
Jesse De Does from the INL gave a brief but rich presentation on the evaluation of lexicon supported OCR and the project’s recent improvements. To evaluate lexica in OCR, the FineReader SDK 10 is used. In short, the software measures OCR with a default included dictionary, and, for each word or fuzzy set, it gives … Continue reading "IMPACT Final Conference – Evaluation of lexicon supported OCR and information retrieval"
IMPACT Final Conference – IBM Adaptive OCR Engine and CONCERT Cooperative Correction
Asaf Tzadok (IBM Haifa Research Lab) showed us IBM’s CONCERT tool which facilitates collaborative OCR correction. CONCERT (Cooperative Engine for the Correction of Extracted Text) works in three steps: character session, word session and page-level session.
IMPACT Final Conference: 1st Keynote: The Strategic Digital Overview
Richard Boulderstone, Director of eStrategy and Programs at the British Library, kicked off the IMPACT Conference this morning with a suitably impactful statement of scope: the British Library, he estimates, has nearly 5 billion physical pages in a 150 million object collection.
Veranstaltungsende
The first part of the OCR workshop in Munich concluded with perspectives of the project IMPACT. The second and more practical orientated part of the OCR workshop can be tracked via the blog of the Munich Centre of Digitisation. Many thanks to the lecturers and sponsors who enabled the exchange of information.
Kollaborative Korrektur
Doris Škarić from the Bavarian State Library reported about collaborative correction of OCR results by volunteers. She presented the IMPACT tool CONCERT (the COllaborative eNgine for the CorREction of Texts) and reported about the findings of a pilot test of the tool at the Bavarian State Library.
Dokumentstrukturerkennung
Günter Mühlberger from the University and Regional Library of Tyrol in Innsbruck presented the Functional Extension Parser (FEP), a tool for the OCR-based structural analysis of printed texts.
Results of OCR Research: IMPACT Demo Day in Munich – Ergebnisse aus der OCR-Forschung: IMPACT Demo Day in München
The Munich DigitiZation Center (MDZ) of the Bavarian State Library invites you to Munich on Tuesday 11 October, for the conference “Results of OCR Research: IMPACT Demo Day”.
Evaluation Framework and Taverna – with Clemens Neudecker
IMPACT Interoperability Framework (IIF) After seeing so many tools introduced during the day, Clemens now considers how you can pull them all together into a usable service for the mass digitisation and OCR of historic text.
Post processing and language technology in OCR
Jesse De Does from INL follows Katrien by introducing a post-processing Text and Error Profiling tool that looks similar to the IBM tool demonstrated by Niall earlier. This tool differs through the use of ‘text profiling’, whereby language and text is analysed on the basis of frequency and logic. It also allows for batch correction … Continue reading "Post processing and language technology in OCR"
The OCR process – Clemens Neudecker & Niall Anderson
IMPACT work with ABBYY IMPACT has been working with both ABBYY and IBM on OCR and Clemens starts with outlining all the steps required during the OCRing process and the extension of the FineReader Engine so it performs better with historical text.
