A large dataset of scientific text reuse in Open-Access publications

被引:3
|
作者
Gienapp, Lukas [1 ]
Kircheis, Wolfgang [1 ,3 ]
Sievers, Bjarne [1 ]
Stein, Benno [2 ]
Potthast, Martin [1 ,3 ]
机构
[1] Univ Leipzig, Text Min & Retrieval Grp, DE-04109 Leipzig, Germany
[2] Bauhaus Univ Weimar, Web Technol & Informat Syst Grp, DE-99423 Weimar, Germany
[3] ScaDS AI, Ctr Scalable Data Analyt & Artificial Intelligence, DE-04105 Leipzig, Germany
关键词
PLAGIARISM;
D O I
10.1038/s41597-022-01908-z
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
We present the Webis-STEREO-21 dataset, a massive collection of Scientific Text Reuse in Open-access publications. It contains 91 million cases of reused text passages found in 4.2 million unique open-access publications. Cases range from overlap of as few as eight words to near-duplicate publications and include a variety of reuse types, ranging from boilerplate text to verbatim copying to quotations and paraphrases. Featuring a high coverage of scientific disciplines and varieties of reuse, as well as comprehensive metadata to contextualize each case, our dataset addresses the most salient shortcomings of previous ones on scientific writing. The Webis-STEREO-21 does not indicate if a reuse case is legitimate or not, as its focus is on the general study of text reuse in science, which is legitimate in the vast majority of cases. It allows for tackling a wide range of research questions from different scientific backgrounds, facilitating both qualitative and quantitative analysis of the phenomenon as well as a first-time grounding on the base rate of text reuse in scientific publications.
引用
收藏
页数:11
相关论文
共 50 条
  • [1] A large dataset of scientific text reuse in Open-Access publications
    Lukas Gienapp
    Wolfgang Kircheis
    Bjarne Sievers
    Benno Stein
    Martin Potthast
    Scientific Data, 10
  • [2] SCIENTIFIC PUBLICATIONS MARKET AND OPEN-ACCESS POLICIES
    Munteanu, R. A.
    Apetroae, M.
    Munteanu, M.
    Iudean, D.
    QUALITY MANAGEMENT IN HIGHER EDUCATION, PROCEEDINGS, 2008, : 553 - 557
  • [3] Why are Latin American countries in the limbo of open-access scientific publications?
    Chavarria, Max
    Lomonte, Bruno
    Gutierrez, Jose Maria
    Moreno, Edgardo
    Perez, Alice L.
    SCIENCE AND PUBLIC POLICY, 2024,
  • [4] Open Access to Scientific Publications
    Beaudouin-Lafon, Michel
    COMMUNICATIONS OF THE ACM, 2010, 53 (02) : 32 - 34
  • [5] Open Access in scientific publications
    Perez Rodrigo, Carmen
    REVISTA ESPANOLA DE NUTRICION COMUNITARIA-SPANISH JOURNAL OF COMMUNITY NUTRITION, 2010, 16 (04): : 203 - 203
  • [6] Open access to scientific publications
    Verma, IM
    MOLECULAR THERAPY, 2004, 10 (04) : 609 - 609
  • [7] Marketing of acquisition for open-access platforms of publications
    Hilse, Stefan
    Depping, Ralf
    BIBLIOTHEK FORSCHUNG UND PRAXIS, 2008, 32 (03) : 334 - 347
  • [8] Open-access Journals - A Scientific Thriller
    Roulet, J. F.
    van Meerbeek, Bart
    JOURNAL OF ADHESIVE DENTISTRY, 2013, 15 (06): : 503 - 504
  • [9] An open-access dataset of emergency department admissions at a large teaching hospital in Iran
    Hassan, Zohreh Bani
    Kazemi, Mohammad Reza
    Jangi, Majid
    Tabesh, Hamed
    DATA IN BRIEF, 2024, 56
  • [10] An Analysis on the Open-Access Publications in "Religious Education" Journal
    Kuscuoglu, Aslihan
    DINBILIMLERI AKADEMIK ARASTIRMA DERGISI-JOURNAL OF ACADEMIC RESEARCH IN RELIGIOUS SCIENCES, 2021, 21 (02): : 1097 - 1128