Category: Discussions
CONCERT crowdsourcing tool – Lotte Wilms & Ewoud Sanders #impactdemo
Manual correction of OCR is expensive and time-consuming. That’s why IBM has developed the CONCERT tool in IMPACT, a crowdsourcing tool with which users can correct OCR results in a very efficient way. The focus is on productivity – the Adaptive OCR engine is trained on the basis of the corrections, which lowers the number of errors.
Experimentele OCR – Lotte Wilms #impactdemo
IMPACT also develops three experimental OCR tools: Wordspotting (developed by NCSR Demokritos) – based on the recognition of complete words Inventory Extraction (developed by the University of Innsbruck) – makes use of character clustering Typewritten OCR (developed by PRImA, Universiteit van Salford) – for typewritten documents
ABBYY Finereader – Lotte Wilms #impactdemo
Lotte Wilms present the various IMPACT improvements that have been incorporated in the new ABBYY Finereader 10 and the ABBYY Recognition Server 3: Better image preprocessing Better segmentation Better recognition quality More languages New optimized processing profiles Native ALTO support
OCR en toepassing bij de KB – Marian Hellema #impactdemo
Marian Hellema is involved with the large-scale newspaper digitisation project at the KB. OCR (Optical Character Recognition) is needed for: Search and retrieval: fulltext search Presentation on a website: highlighting search terms or text-only presentation
IMPACT Beeldverbetering – Lotte Wilms #impactdemo
Lotte Wilms shows various ways to enhance an image with IMPACT tools after scanning to improve OCR results.
Kennisbank Digitalisering – Lotte Wilms #impactdemo
Various library partners in IMPACT (BnF, ONB, BL, KB, DNB, UGOE and BSB) are working together on the Decision Support Tools to make sure the broader public has access to the correct information on digitisation. These libraries all face similar problems, such as complex layout, historical and gothic fonts or damaged material.
The Question & Answer Session
Better late than never, a short summary of the question and answer session. NOT a literal transcript. Unfortunately, Sven Schlarb couldn’t take part in the session.
Use of digitised and OCRed text collections by end users
Geneviève Cron of the Bibliotheque Nationale de France (BnF) begins by discussing the BNF’s digital library: Gallica. A million documents digitised since 1992, with OCR as standard since 2005. OCR accuracy for newspapers is 98% on word level, but results are much more varied – from 60% up. For books, the average accuracy lies at … Continue reading "Use of digitised and OCRed text collections by end users"
A gentle introduction to lexicon building and application
Katrien Depuydt of the Institute for Dutch Lexicology (INL) begins with the distinction between a lexicon and an electronic dictionary. Dictionaries are primarily for human use and organised so that each entry is an item in itself. Lexica on the other hand are primarily for computation, and used for linguistic annotation, enhanced retrieval (like tracking … Continue reading "A gentle introduction to lexicon building and application"
The Functional Extension Parser – a rule-based system for flexible structural analysis
Lukas Gander of Universitäts- und Landesbibliothek Tirol (University and Regional Library Tyrol) outlines the concept behind the Functional Extension Parser: using an OCR engine’s output to create a structural map of a page or volume. OCR engines capture much more information than simple text: for instance, they contain information about text type and position. The … Continue reading "The Functional Extension Parser – a rule-based system for flexible structural analysis"
