A Piggyback System for Joint Entity Mention Detection and Linking in Web Queries

被引:19
作者
Cornolti, Marco [1 ]
Ferragina, Paolo [1 ]
Ciaramita, Massimiliano [2 ]
Rued, Stefan [3 ]
Schuetze, Hinrich [3 ]
机构
[1] Univ Pisa, Pisa, Italy
[2] Google, Zurich, Switzerland
[3] Univ Munich, Munich, Germany
来源
PROCEEDINGS OF THE 25TH INTERNATIONAL CONFERENCE ON WORLD WIDE WEB (WWW'16) | 2016年
基金
欧盟地平线“2020”;
关键词
Entity linking; query annotation; ERD; piggyback;
D O I
10.1145/2872427.2883061
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
In this paper we study the problem of linking open-domain web-search queries towards entities drawn from the full entity inventory of Wikipedia articles. We introduce SMAPH2, a second-order approach that, by piggybacking on a web search engine, alleviates the noise and irregularities that characterize the language of queries and puts queries in a larger context in which it is easier to make sense of them. The key algorithmic idea underlying SMAPH-2 is to first discover a candidate set of entities and then link-back those entities to their mentions occurring in the input query. This allows us to confine the possible concepts pertinent to the query to only the ones really mentioned in it. The link-back is implemented via a collective disambiguation step based upon a supervised ranking model that makes one joint prediction for the annotation of the complete query optimizing directly the F1 measure. We evaluate both known features, such as word embeddings and semantic relatedness among entities, and several novel features such as an approximate distance between mentions and entities (which can handle spelling errors). We demonstrate that SMAPH-2 achieves state-of-the-art performance on the ERD@SIGIR2014 benchmark. We also publish GERDAQ (General Entity Recognition, Disambiguation and Annotation in Queries), a novel, public dataset built specifically for web-query entity linking via a crowdsourcing effort. SMAPH-2 outperforms the benchmarks by comparable margins also on GERDAQ.
引用
收藏
页码:567 / 578
页数:12
相关论文
共 43 条
[1]  
Alasiry A, 2012, SIGIR 2012: PROCEEDINGS OF THE 35TH INTERNATIONAL ACM SIGIR CONFERENCE ON RESEARCH AND DEVELOPMENT IN INFORMATION RETRIEVAL, P1049, DOI 10.1145/2348283.2348463
[2]  
[Anonymous], 2008, Proceedings of the 17th ACM conference on Information and knowledge management
[3]  
[Anonymous], 2011, P 2011 C EMPIRICAL M, DOI DOI 10.3115/V1/D11-1072
[4]  
[Anonymous], 2014, EMNLP
[5]  
[Anonymous], 2011, P 49 ANN M ASS COMP
[6]  
Bendersky M., 2011, ACL, P102
[7]   Fast and Space-Efficient Entity Linking in Queries [J].
Blanco, Roi ;
Ottaviano, Giuseppe ;
Meij, Edgar .
WSDM'15: PROCEEDINGS OF THE EIGHTH ACM INTERNATIONAL CONFERENCE ON WEB SEARCH AND DATA MINING, 2015, :179-188
[8]  
Bordino I., 2013, WSDM, P275, DOI DOI 10.1145/2433396.2433433
[9]  
Carmel D., 2014, SIGIR FORUM
[10]   LIBSVM: A Library for Support Vector Machines [J].
Chang, Chih-Chung ;
Lin, Chih-Jen .
ACM TRANSACTIONS ON INTELLIGENT SYSTEMS AND TECHNOLOGY, 2011, 2 (03)