Wikipedia citations: A comprehensive data set of citations with identifiers extracted from English Wikipedia

被引:17
作者
Singh, Harshdeep [1 ]
West, Robert [1 ]
Colavizza, Giovanni [2 ]
机构
[1] Ecole Polytech Fed Lausanne, Data Sci Lab, Lausanne, Switzerland
[2] Univ Amsterdam, Inst Log Language & Computat, Amsterdam, Netherlands
来源
QUANTITATIVE SCIENCE STUDIES | 2021年 / 2卷 / 01期
关键词
citations; data; data set; Wikipedia; KNOWLEDGE; SCIENCE;
D O I
10.1162/qss_a_00105
中图分类号
G25 [图书馆学、图书馆事业]; G35 [情报学、情报工作];
学科分类号
1205 ; 120501 ;
摘要
Wikipedia's content is based on reliable and published sources. To this date, relatively little is known about what sources Wikipedia relies on, in part because extracting citations and identifying cited sources is challenging. To close this gap, we release Wikipedia Citations, a comprehensive data set of citations extracted from Wikipedia. We extracted29.3 million citations from 6.1 million English Wikipedia articles as of May 2020, and classified as being books, journal articles, or Web content. We were thus able to extract 4.0 million citations to scholarly publications with known identifiers-including DOI, PMC, PMID, and ISBN-and further equip an extra 261 thousand citations with DOIs from Crossref. As a result, we find that 6.7% of Wikipedia articles cite at least one journal article with an associated DOI, and that Wikipedia cites just 2% of all articles with a DOI currently indexed in the Web of Science. We release our code to allow the community to extend upon our work and update the data set in the future.
引用
收藏
页码:1 / 19
页数:19
相关论文
共 57 条
[1]  
[Anonymous], 2015, ACS SYM SER
[2]  
[Anonymous], **DATA OBJECT**, DOI DOI 10.5281/ZENODO.3940692
[3]  
[Anonymous], **DATA OBJECT**, DOI DOI 10.6084/M9.FIGSHARE.1299540
[4]   Science through Wikipedia: A novel representation of open knowledge through co-citation networks [J].
Arroyo-Machado, Wenceslao ;
Torres-Salinas, Daniel ;
Herrera-Viedma, Enrique ;
Romero-Frias, Esteban .
PLOS ONE, 2020, 15 (02)
[5]   A Graph-Structured Dataset for Wikipedia Research [J].
Aspert, Nicolas ;
Miz, Volodymyr ;
Ricaud, Benjamin ;
Vandergheynst, Pierre .
COMPANION OF THE WORLD WIDE WEB CONFERENCE (WWW 2019 ), 2019, :1188-1193
[6]   Web of Science as a data source for research on scientific and scholarly activity [J].
Birkle, Caroline ;
Pendlebury, David A. ;
Schnell, Joshua ;
Adams, Jonathan .
QUANTITATIVE SCIENCE STUDIES, 2020, 1 (01) :363-376
[7]  
Bojanowski Piotr, 2017, Transactions of the Association for Computational Linguistics, V5, P135, DOI DOI 10.1162/TACL_A_00051
[8]  
Borner Katy, 2010, ATLAS SCI VISUALIZIN
[9]  
CHEN CC, 2012, P 8 ANN INT S WIK OP, DOI DOI 10.1145/2462932.2462943
[10]  
Chen CM, 2017, J DATA INFO SCI, V2, P1, DOI 10.1515/jdis-2017-0006