Cost estimation for 100% postcorrection


Which will be the cost estimation for a project pursuing 100% post correction of OCR in the full text offered to the users (based on an average book of 200-300 pages)? An example is the Gutenberg project (http://www.gutenberg.org/). Has your institution experience on this service?

What’s in a word?


Word error rate is often used as a measure of OCR accuracy. Although words are the relevant unit in information retrieval, the definition of word in the context of OCR is not as simple as it might appear at first glance. This post compiles some of my ideas on the relation between words and characters.

Punctuation in OCR evaluation


While evaluating OCR, some users demand that diacritics should not be considered an error since they are often neglected for information retrieval (search engines). This request raises some doubts about the expected behavior. What do you think about the following issues?

Evaluation Framework and Taverna – with Clemens Neudecker


IMPACT Interoperability Framework (IIF) After seeing so many tools introduced during the day, Clemens now considers how you can pull them all together into a usable service for the mass digitisation and OCR of historic text.

Experimentele OCR – Lotte Wilms #impactdemo


IMPACT also develops three experimental OCR tools: Wordspotting (developed by NCSR Demokritos) – based on the recognition of complete words Inventory Extraction (developed by the University of Innsbruck) – makes use of character clustering Typewritten OCR (developed by PRImA, Universiteit van Salford) – for typewritten documents