Studying the XML Web: Gathering Statistics from an XML Sample

被引:0
作者
Denilson Barbosa
Laurent Mignet
Pierangelo Veltri
机构
[1] University of Toronto,Department of Computer Science
[2] IBM India Research Laboratory,Department of Experimental and Clinical Medicine
[3] Magna Graecia University of Catanzaro,undefined
来源
World Wide Web | 2005年 / 8卷
关键词
World Wide Web; XML; XML web; XML Documents; XML processing tools;
D O I
暂无
中图分类号
学科分类号
摘要
XML has emerged as the language for exchanging data on the web and has attracted considerable interest both in industry and in academia. Nevertheless, to date, little is known about the XML documents published on the web. This paper presents a comprehensive analysis of a sample of about 200,000 XML documents on the web, and is the first study of its kind. We study the distribution of XML documents across the web in several ways; moreover, we provided a detailed characterization of the structure of real XML documents. Our results provide valuable input to the design of algorithms, tools and systems that use XML in one form or another.
引用
收藏
页码:413 / 438
页数:25
相关论文
共 18 条
[1]  
Fiebig T.(2002)Anatomy of a native XML base management system VLDB Journal 11 292-314
[2]  
Helmer S.(2002)TIMBER: A native XML database VLDB Journal 11 274-291
[3]  
Kanne C.(undefined)undefined undefined undefined undefined-undefined
[4]  
Moerkotte G.(undefined)undefined undefined undefined undefined-undefined
[5]  
Neumann J.(undefined)undefined undefined undefined undefined-undefined
[6]  
Schiele R.(undefined)undefined undefined undefined undefined-undefined
[7]  
Westmann T.(undefined)undefined undefined undefined undefined-undefined
[8]  
Jagadish H. V.(undefined)undefined undefined undefined undefined-undefined
[9]  
Al-Khalifa S.(undefined)undefined undefined undefined undefined-undefined
[10]  
Chapman A.(undefined)undefined undefined undefined undefined-undefined