Digital weight watching: reconstruction of scanned documents

Open Access
Authors
Publication date 2011
Journal International Journal on Document Analysis and Recognition
Volume | Issue number 14 | 2
Pages (from-to) 229-239
Organisations
  • Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract
A web portal providing access to over 250.000 scanned and OCRed cultural heritage documents is analyzed. The collection consists of the complete Dutch Hansard from 1917 to 1995. Each document consists of facsimile images of the original pages plus hidden OCRed text. The inclusion of images yields large file sizes of which less than 2% is the actual text. The search user interface of the portal provides poor ranking and not very informative document summaries (snippets). Thus, users are instrumental in weeding out non-relevant results. For that, they must assess the complete documents. This is a time-consuming and frustrating process because of long download and processing times of the large files. Instead of using the scanned images for relevance assessment, we propose to use a reconstruction of the original document from a purely semantic representation. Evaluation on the Dutch dataset shows that these reconstructions become two orders of magnitude smaller and still resemble the original to a high degree. In addition, they are easier to speed-read and evaluate for relevance, due to added hyperlinks and a presentation optimized for reading from a terminal. We describe the reconstruction process and evaluate the costs, the benefits, and the quality.
Document type Article
Language English
Published at https://doi.org/10.1007/s10032-010-0135-3
Downloads
361287.pdf (Final published version)
Permalink to this page
Back