Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond

被引：12

作者：

Karagiannis, Georgios ^{[1
]}

Trummer, Immanuel ^{[1
]}

Jo, Saehan ^{[1
]}

Khandelwal, Shubham ^{[1
]}

Wang, Xuezhi ^{[2
]}

Yu, Cong ^{[2
]}

机构：

[1] Cornell Univ, Ithaca, NY 14853 USA

[2] Google, New York, NY USA

来源：

PROCEEDINGS OF THE VLDB ENDOWMENT | 2019年 / 13卷 / 04期

关键词：

D O I：

10.14778/3372716.3372727

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

We introduce the problem of anti-knowledge mining. Our goal is to create an "anti-knowledge base" that contains factual mistakes. The resulting data can be used for analysis, training, and benchmarking in the research domain of automated fact checking. Prior data sets feature manually generated fact checks of famous misclaims. Instead, we focus on the long tail of factual mistakes made by Web authors, ranging from erroneous sports results to incorrect capitals. We mine mistakes automatically, by an unsupervised approach, from Wikipedia updates that correct factual mistakes. Identifying such updates (only a small fraction of the total number of updates) is one of the primary challenges. We mine anti-knowledge by a multi-step pipeline. First, we filter out candidate updates via several simple heuristics. Next, we correlate Wikipedia updates with other statements made on the Web. Using claim occurrence frequencies as input to a probabilistic model, we infer the likelihood of corrections via an iterative expectation-maximization approach. Finally, we extract mistakes in the form of subject-predicate-object triples and rank them according to several criteria. Our end result is a data set containing over 110,000 ranked mistakes with a precision of 85% in the top 1% and a precision of over 60% in the top 25%. We demonstrate that baselines achieve significantly lower precision. Also, we exploit our data to verify several hypothesis on why users make mistakes. We finally show that the AKB can be used to find mistakes on the entire Web.

引用

页码：561 / 573

页数：13

共 32 条

[1]

Angeli G, 2015, PROCEEDINGS OF THE 53RD ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS AND THE 7TH INTERNATIONAL JOINT CONFERENCE ON NATURAL LANGUAGE PROCESSING, VOL 1, P344

[2]

[Anonymous], 2007, P 33 INT C VER LARG

[3]

[Anonymous], 2018, VLDB, DOI DOI 10.1145/3177732.3177739

[4]

Aydin BI, 2014, AAAI CONF ARTIF INTE, P2946

[5] Navigating the Maze of Wikidata Query Logs [J].

Bonifati, Angela ;

Martens, Wim ;

Timm, Thomas .

WEB CONFERENCE 2019: PROCEEDINGS OF THE WORLD WIDE WEB CONFERENCE (WWW 2019), 2019, :127-138

[6]

Dawid A. P., 1979, J ROYAL STAT SOC SER, V28, P20

[7]

De Sa C. M., 2015, NIPS, P3097

[8]

Devlin J, 2019, 2019 CONFERENCE OF THE NORTH AMERICAN CHAPTER OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS: HUMAN LANGUAGE TECHNOLOGIES (NAACL HLT 2019), VOL. 1, P4171

[9]

Etzioni O., 2004, P 13 INT C WORLD WID, P100, DOI [DOI 10.1145/988672.988687, 10.1145/988672.988687]

[10]

Fader A., 2011, Identifying Relations for Open Information Extraction

← 1 2 3 4 →