A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration

被引:217
作者
Zhao, Bo [1 ]
Rubinstein, Benjamin I. P. [2 ]
Gemmell, Jim [2 ]
Han, Jiawei [1 ]
机构
[1] Univ Illinois, Dept Comp Sci, 1304 W Springfield Ave, Urbana, IL 61801 USA
[2] Microsoft Res, Mountain View, CA USA
来源
PROCEEDINGS OF THE VLDB ENDOWMENT | 2012年 / 5卷 / 06期
关键词
Bayesian networks - Inference engines;
D O I
10.14778/2168651.2168656
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
In practical data integration systems, it is common for the data sources being integrated to provide conflicting information about the same entity. Consequently, a major challenge for data integration is to derive the most complete and accurate integrated records from diverse and sometimes conflicting sources. We term this challenge the truth finding problem. We observe that some sources are generally more reliable than others, and therefore a good model of source quality is the key to solving the truth finding problem. In this work, we propose a probabilistic graphical model that can automatically infer true records and source quality without any supervision. In contrast to previous methods, our principled approach leverages a generative process of two types of errors (false positive and false negative) by modeling two different aspects of source quality. In so doing, ours is also the first approach designed to merge multi-valued attribute types. Our method is scalable, due to an efficient sampling-based inference algorithm that needs very few iterations in practice and enjoys linear time complexity, with an even faster incremental variant. Experiments on two real world datasets show that our new method outperforms existing state-of-the-art approaches to the truth finding problem.
引用
收藏
页码:550 / 561
页数:12
相关论文
共 15 条
[1]  
[Anonymous], 2010, P WSDM 2010 3 INT C
[2]  
[Anonymous], 2009, PROC VLDB ENDOW, DOI DOI 10.14778/1687627.1687691
[3]  
Arenas M., 1999, Proceedings of the Eighteenth ACM SIGMOD-SIGACT-SIGART Symposium on Principles of Database Systems, P68, DOI 10.1145/303976.303983
[4]  
Balakrishnan R., 2011, P 20 INT C WORLD WID, V1, P227
[5]  
Blanco L, 2010, LECT NOTES COMPUT SC, V6051, P83, DOI 10.1007/978-3-642-13094-6_8
[6]  
Dong X L, 2009, PROC VLDB ENDOW, P550, DOI DOI 10.14778/1687627.1687690
[7]  
Florescu D, 1997, PROCEEDINGS OF THE TWENTY-THIRD INTERNATIONAL CONFERENCE ON VERY LARGE DATABASES, P216
[8]  
Kasneci G., 2011, P 4 ACM INT C WEB SE, P465
[9]   Authoritative sources in a hyperlinked environment [J].
Kleinberg, JM .
JOURNAL OF THE ACM, 1999, 46 (05) :604-632
[10]  
Pasternack J, 2010, P 23 INT C COMP LING, P877