The quality of the XML web

Authors	S. Grijzenhout M. Marx
Publication date	2011
Book title	CIKM'11
Book subtitle	proceedings of the 2011 ACM International Conference on Information and Knowledge Management : October 24-28, 2011, Glasgow, Scotland
ISBN (electronic)	9781450307178
Event	2011 ACM International Conference on Information and Knowledge Management
Pages (from-to)	1719-1724
Publisher	New York, NY: Association for Computing Machinery
Organisations	Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract	We collect evidence to answer the following question: Is the quality of the XML documents found on the web sufficient to apply XML technology like XQuery, XPath and XSLT? XML collections from the web have been previously studied statistically, but no detailed information about the quality of the XML documents on the web is available to date. We address this shortcoming in this study. We gathered 180K XML documents from the web. Their quality is surprisingly good; 85.4% is well-formed and 99.5% of all specified encodings is correct. Validity needs serious attention. Only 25% of all files contain a reference to a DTD or XSD, of which just one third is actually valid. Errors are studied in detail. Automatic error repair seems promising. Our study is well documented and easily repeatable. This paves the way for a periodic quality assessment of the XML web.
Document type	Conference contribution
Language	English
Published at	https://doi.org/10.1145/2063576.2063824 (Final published version)
Permalink to this page

Back

UvA-DARE