Digital sustainable publication of legacy parliamentary proceedings

Authors
Publication date 2010
Book title Proceedings 11th International Digital Government Research Conference (dg.o 2010)
ISBN
  • 9781450300704
Event 11th Annual International Digital Government Research Conference on Public Administration Online: Challenges and Opportunities (dg.o '10), Puebla, Mexico
Pages (from-to) 99-104
Publisher Digital Government Society of North America
Organisations
  • Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract
We address the problem of publishing parliamentary proceedings in a digital sustainable manner. We give an extensive requirements analysis, and based on that propose a uniform XML format. We evaluated our approach by collecting and automatically processing proceedings from six parliaments spanning almost 200 years in total. Most of this data is real legacy data consisting of scanned and OCRed documents. The approach scales very well and produces high quality data.
All documents are transformed into UTF-8 encoded XML files with extensive metadata in Dublin Core standard. The text itself is divided into pages which are divided into paragraphs. Every document, page and paragraph has a unique URN which resolves to a web page. Every page element in the XML files is connected to a facsimile image of that page in PDF or JPEG format. We created a viewer in which both versions can be inspected simultaneously. A search-engine for the complete collection is available online.
Document type Conference contribution
Language English
Published at http://portal.acm.org/citation.cfm?id=1809895
Permalink to this page
Back