Combining Tag and Value Similarity for Data Extraction and Alignment

被引:17
作者
Su, Weifeng [1 ]
Wang, Jiying [2 ]
Lochovsky, Frederick H. [3 ]
Liu, Yi [4 ]
机构
[1] BNU HKBU United Int Coll, Comp Sci & Technol Program, Tangjiawan, Zhuhai, Peoples R China
[2] City Univ Hong Kong, Dept Comp Sci, Kowloon, Hong Kong, Peoples R China
[3] Hong Kong Univ Sci & Technol, Dept Comp Sci & Engn, Kowloon, Hong Kong, Peoples R China
[4] Tsinghua Natl Lab Informat Sci & Technol, Div Technol Innovat & Dev, Ctr Speech & Language Technol, Beijing, Peoples R China
基金
中国国家自然科学基金;
关键词
Data extraction; automatic wrapper generation; data record alignment; information integration; WRAPPER INDUCTION; WEB;
D O I
10.1109/TKDE.2011.66
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Web databases generate query result pages based on a user's query. Automatically extracting the data from these query result pages is very important for many applications, such as data integration, which need to cooperate with multiple web databases. We present a novel data extraction and alignment method called CTVS that combines both tag and value similarity. CTVS automatically extracts data from query result pages by first identifying and segmenting the query result records (QRRs) in the query result pages and then aligning the segmented QRRs into a table, in which the data values from the same attribute are put into the same column. Specifically, we propose new techniques to handle the case when the QRRs are not contiguous, which may be due to the presence of auxiliary information, such as a comment, recommendation or advertisement, and for handling any nested structure that may exist in the QRRs. We also design a new record alignment algorithm that aligns the attributes in a record, first pairwise and then holistically, by combining the tag and data value similarity information. Experimental results show that CTVS achieves high precision and outperforms existing state-of-the-art data extraction methods.
引用
收藏
页码:1186 / 1200
页数:15
相关论文
共 31 条
  • [1] [Anonymous], P 3 INT C WEB INF SY
  • [2] [Anonymous], 1997, ACM SIGACT NEWS
  • [3] Arasu A., 2003, P 2003 ACM SIGMOD IN, P337, DOI DOI 10.1145/872757.872799
  • [4] Baeza-Yates R., 1989, ACM SPECIAL INTEREST, V23, P34
  • [5] Baumgartner R., 2001, Proceedings of the 27th International Conference on Very Large Data Bases, P119
  • [6] Bergman MichaelK., 2001, DEEP WEB SURFACING H
  • [7] The complexity of multiple sequence alignment with SP-score that is a metric
    Bonizzoni, P
    Della Vedova, G
    [J]. THEORETICAL COMPUTER SCIENCE, 2001, 259 (1-2) : 63 - 79
  • [8] A fully automated object extraction system for the World Wide Web
    Buttler, D
    Liu, L
    Pu, C
    [J]. 21ST INTERNATIONAL CONFERENCE ON DISTRIBUTED COMPUTING SYSTEMS, PROCEEDINGS, 2001, : 361 - 370
  • [9] Self-pumped and mutually pumped phase conjugation in pentagon-shaped BaTiO3 crystal with plus c-face incident geometry
    Chang, CC
    Chen, TC
    Hu, GW
    Yau, HF
    Ye, PX
    [J]. PHOTOREFRACTIVE EFFECTS, MATERIALS AND DEVICES, PROCEEDINGS, 2001, 62 : 681 - 681
  • [10] Chang KCC, 2004, SIGMOD REC, V33, P61, DOI 10.1145/1031570.1031584