Large-scale linked data integration using probabilistic reasoning and crowdsourcing

被引:0
作者
Gianluca Demartini
Djellel Eddine Difallah
Philippe Cudré-Mauroux
机构
[1] University of Fribourg,eXascale Infolab
来源
The VLDB Journal | 2013年 / 22卷
关键词
Instance matching; Entity linking; Data integration; Crowdsourcing; Probabilistic reasoning;
D O I
暂无
中图分类号
学科分类号
摘要
We tackle the problems of semiautomatically matching linked data sets and of linking large collections of Web pages to linked data. Our system, ZenCrowd, (1) uses a three-stage blocking technique in order to obtain the best possible instance matches while minimizing both computational complexity and latency, and (2) identifies entities from natural language text using state-of-the-art techniques and automatically connects them to the linked open data cloud. First, we use structured inverted indices to quickly find potential candidate results from entities that have been indexed in our system. Our system then analyzes the candidate matches and refines them whenever deemed necessary using computationally more expensive queries on a graph database. Finally, we resort to human computation by dynamically generating crowdsourcing tasks in case the algorithmic components fail to come up with convincing results. We integrate all results from the inverted indices, from the graph database and from the crowd using a probabilistic framework in order to make sensible decisions about candidate matches and to identify unreliable human workers. In the following, we give an overview of the architecture of our system and describe in detail our novel three-stage blocking technique and our probabilistic decision framework. We also report on a series of experimental results on a standard data set, showing that our system can achieve a 95 % average accuracy on instance matching (as compared to the initial 88 % average accuracy of the purely automatic baseline) while drastically limiting the amount of work performed by the crowd. The experimental evaluation of our system on the entity linking task shows an average relative improvement of 14 % over our best automatic approach.
引用
收藏
页码:665 / 687
页数:22
相关论文
共 25 条
[1]  
Christen P(2012)A survey of indexing techniques for scalable record linkage and deduplication IEEE Trans. Knowl. Data Eng. 24 1537-1555
[2]  
Feng A(2011)CrowdDB: Query Processing with the VLDB Crowd PVLDB 4 1387-1390
[3]  
Franklin MJ(1989)Advances in record-linkage methodology as applied to matching the 1985 census of Tampa Florida. J. Am. Stat. Assoc. 84 414-420
[4]  
Kossmann D(1996)Entity identification in database integration Inform. Sci. 89 1-38
[5]  
Kraska T(2011)Human-powered sorts and joins PVLDB 5 13-24
[6]  
Madden S(2008)Designing games with a purpose Commun. ACM 51 58-67
[7]  
Ramesh S(2012)CrowdER: crowdsourcing entity resolution PVLDB 5 1483-1494
[8]  
Wang A(undefined)undefined undefined undefined undefined-undefined
[9]  
Xin R(undefined)undefined undefined undefined undefined-undefined
[10]  
Jaro M(undefined)undefined undefined undefined undefined-undefined