Digital weight watching: reconstruction of scanned documents

M. Marx; T. Gielissen

doi:https://doi.org/10.1007/s10032-010-0135-3

Digital weight watching: reconstruction of scanned documents

Authors	M. Marx T. Gielissen
Publication date	2011
Journal	International Journal on Document Analysis and Recognition
Volume \| Issue number	14 \| 2
Pages (from-to)	229-239
Organisations	Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract	A web portal providing access to over 250.000 scanned and OCRed cultural heritage documents is analyzed. The collection consists of the complete Dutch Hansard from 1917 to 1995. Each document consists of facsimile images of the original pages plus hidden OCRed text. The inclusion of images yields large file sizes of which less than 2% is the actual text. The search user interface of the portal provides poor ranking and not very informative document summaries (snippets). Thus, users are instrumental in weeding out non-relevant results. For that, they must assess the complete documents. This is a time-consuming and frustrating process because of long download and processing times of the large files. Instead of using the scanned images for relevance assessment, we propose to use a reconstruction of the original document from a purely semantic representation. Evaluation on the Dutch dataset shows that these reconstructions become two orders of magnitude smaller and still resemble the original to a high degree. In addition, they are easier to speed-read and evaluate for relevance, due to added hyperlinks and a presentation optimized for reading from a terminal. We describe the reconstruction process and evaluate the costs, the benefits, and the quality.
Document type	Article
Language	English
Published at	https://doi.org/10.1007/s10032-010-0135-3 (Final published version)
Downloads	361287.pdf (Final published version)
Permalink to this page

Back

UvA-DARE

Digital Academic Repository

Digital weight watching: reconstruction of scanned documents