Author: job.vandoeselaar@ivdnt.org
End of the workshop!
The workshop wrapped up with a testing Q&A from delegates about the potential uses and availability of the IMPACT Tools. A full account of that session will follow next week, along with videos of each of the presentations.
Use of digitised and OCRed text collections by end users
Geneviève Cron of the Bibliotheque Nationale de France (BnF) begins by discussing the BNF’s digital library: Gallica. A million documents digitised since 1992, with OCR as standard since 2005. OCR accuracy for newspapers is 98% on word level, but results are much more varied – from 60% up. For books, the average accuracy lies at … Continue reading "Use of digitised and OCRed text collections by end users"
A gentle introduction to lexicon building and application
Katrien Depuydt of the Institute for Dutch Lexicology (INL) begins with the distinction between a lexicon and an electronic dictionary. Dictionaries are primarily for human use and organised so that each entry is an item in itself. Lexica on the other hand are primarily for computation, and used for linguistic annotation, enhanced retrieval (like tracking … Continue reading "A gentle introduction to lexicon building and application"
Tools for Document Image Analysis
[slideshare id=4138338&doc=bratislavaws-pratikakis-ncsr-imageanalysistoolsnew-100518090539-phpapp02] http://vimeo.com/11833935 Ioannis Pratikakis of the NCSR – National Center for Scientific Research – “Demokritos” now provides a live demo of various image pre-processing tools. By combining various familiar algorithms used to binarise images, an operator can get a good visual idea of which type of binarisation (or combination) will produce the best OCR … Continue reading "Tools for Document Image Analysis"
The Functional Extension Parser – a rule-based system for flexible structural analysis
Lukas Gander of Universitäts- und Landesbibliothek Tirol (University and Regional Library Tyrol) outlines the concept behind the Functional Extension Parser: using an OCR engine’s output to create a structural map of a page or volume. OCR engines capture much more information than simple text: for instance, they contain information about text type and position. The … Continue reading "The Functional Extension Parser – a rule-based system for flexible structural analysis"
Working in CONCERT – public participation in mass digitisation
Next on stage is Niall Anderson from the British Library, talking about public participation in mass digitisation, and “why we think that’s a good idea”. A 2008 Conference of European Libraries survey estimated that that there were now some 8 million digitised text-based items in existence, proof that we live in an era of effective … Continue reading "Working in CONCERT – public participation in mass digitisation"
The challenges of historical materials and an overview on the technical solutions in IMPACT
Sven Schlarb of the Österreichische Nationalbibliothek (Austrian National Library) now talks about the challenges of text digitisation for OCR and the solutions IMPACT has devised to deal with them. Having outlined the individual tools and the partners responsible for their development, he talks in detail about the ideal IMPACT workflow in which they can all … Continue reading "The challenges of historical materials and an overview on the technical solutions in IMPACT"
Overview of the IMPACT Project
Aly Conteh returns to the stage to explain the context in which IMPACT sees itself operating. He explains first of all The British Library’s particular interest in the project: even material that had already been digitised had not always been digitised in such a way as to be accessible, except as an illustrative image. The … Continue reading "Overview of the IMPACT Project"
OCR in libraries – some practical remarks
Next, Günter Mühlberger of Universitäts- und Landesbibliothek Tirol (University and Regional Library Tyrol) talks about the situation with regard to full-text digitisation in Europe and elsewhere. He notes that in Europe, OCR has not always been a standard part of the workflow, so there is much legacy digital material that has never been OCR’d – … Continue reading "OCR in libraries – some practical remarks"
Optical Character Recognition – introduction and overview
Michael Fuchs of Abbyy starts by explaining the work of his company within IMPACT: Abbyy provides the other partners with access to its FineReader SDK, and uses the experiments and digital text material within IMPACT to hone and test its own products and technology. He identifies Fraktur/Gothic script as being a particular difficulty for state … Continue reading "Optical Character Recognition – introduction and overview"
