Succeed survey on content/tools licensing and innovative usages


The Succeed project is undertaking a survey on the use of licenses in the field of digitisation and on innovative usages of digitised content. The aim of this survey is to gather information about current practices for licensing data, metadata and tools and on new trends in the exploitation of digitised content. This information will help define … Continue reading "Succeed survey on content/tools licensing and innovative usages"

Working together to improve text digitisation techniques


2nd Succeed hackathon at the University of Alicante Is there anyone out there still thinking that a hackathon is a malicious break-in? Far from it. It is the best way for developers and researchers to get together and work on new tools and innovations. The2nd developers workshop / hackathon organised on 10-11 April by the Succeed Project was a case … Continue reading "Working together to improve text digitisation techniques"

How to maximise usage of digital collections


(Reblogged from Research in KB blog: http://researchkb.wordpress.com/2014/04/13/how-to-maximise-usage-of-digital-collections/) Libraries want to understand the researchers who use their digital collections and researchers want to understand the nature of these collections better. The seminar ‘Mining digital repositories’ brought them together at the Dutch Koninklijke Bibliotheek (KB) on 10-11 April, 2014,

Cost estimation for 100% postcorrection


Which will be the cost estimation for a project pursuing 100% post correction of OCR in the full text offered to the users (based on an average book of 200-300 pages)? An example is the Gutenberg project (http://www.gutenberg.org/). Has your institution experience on this service?

What’s in a word?


Word error rate is often used as a measure of OCR accuracy. Although words are the relevant unit in information retrieval, the definition of word in the context of OCR is not as simple as it might appear at first glance. This post compiles some of my ideas on the relation between words and characters.

Punctuation in OCR evaluation


While evaluating OCR, some users demand that diacritics should not be considered an error since they are often neglected for information retrieval (search engines). This request raises some doubts about the expected behavior. What do you think about the following issues?

Succeed training materials


In the framework of the Succeed project, a set of tools for digitisation has been selected for recommendation. In concert with the library partners a subset of these tools is selected that will be actually deployed in pilot projects.

iConference 2014 Workshop. Berlin, March 4th, 2014.


Digital Collection Contexts: Intellectual and Organizational Functions at Scale http://pro.europeana.eu/web/network/europeana-tech/-/wiki/Main/iConference2014 What are digital collections and how are they to be accessed? This workshop organized by Europeana Pro addresses also the information needs of scholars and aspects of international interoperability.

Faster, smarter rand richer 2014. Reshaping the library catalogue. International conference.


Traditional library catalogues serve to locate printed publications. The advance of electronic publications and digitisation is changing the way users want to locate and access publications. This brings into question the role of the catalogue.

How can I select the best OCR engine for my digitisation requirements?


I guess there is no “best” OCR engine independently of the purpose but, is there any protocol to compare their performance? Perhaps, the Centre of Competence could provide some sample sets and evaluation measures on those sample sets. If they provide sufficient coverage, probably my requirements are similar to one of the subsets and it … Continue reading "How can I select the best OCR engine for my digitisation requirements?"