A large dataset of scientific text reuse in Open-Access publications

被引:3
|
作者
Gienapp, Lukas [1 ]
Kircheis, Wolfgang [1 ,3 ]
Sievers, Bjarne [1 ]
Stein, Benno [2 ]
Potthast, Martin [1 ,3 ]
机构
[1] Univ Leipzig, Text Min & Retrieval Grp, DE-04109 Leipzig, Germany
[2] Bauhaus Univ Weimar, Web Technol & Informat Syst Grp, DE-99423 Weimar, Germany
[3] ScaDS AI, Ctr Scalable Data Analyt & Artificial Intelligence, DE-04105 Leipzig, Germany
关键词
PLAGIARISM;
D O I
10.1038/s41597-022-01908-z
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
We present the Webis-STEREO-21 dataset, a massive collection of Scientific Text Reuse in Open-access publications. It contains 91 million cases of reused text passages found in 4.2 million unique open-access publications. Cases range from overlap of as few as eight words to near-duplicate publications and include a variety of reuse types, ranging from boilerplate text to verbatim copying to quotations and paraphrases. Featuring a high coverage of scientific disciplines and varieties of reuse, as well as comprehensive metadata to contextualize each case, our dataset addresses the most salient shortcomings of previous ones on scientific writing. The Webis-STEREO-21 does not indicate if a reuse case is legitimate or not, as its focus is on the general study of text reuse in science, which is legitimate in the vast majority of cases. It allows for tackling a wide range of research questions from different scientific backgrounds, facilitating both qualitative and quantitative analysis of the phenomenon as well as a first-time grounding on the base rate of text reuse in scientific publications.
引用
收藏
页数:11
相关论文
共 50 条
  • [21] PatCID: an open-access dataset of chemical structures in patent documents
    Morin, Lucas
    Weber, Valery
    Meijer, Gerhard Ingmar
    Yu, Fisher
    Staar, Peter W. J.
    NATURE COMMUNICATIONS, 2024, 15 (01)
  • [22] Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning
    Singh, Shivalika
    Vargus, Freddie
    D'souza, Daniel
    Karlsson, Borje F.
    Mahendiran, Abinaya
    Ko, Wei-Yin
    Shandilyal, Herumb
    Patel, Jay
    Mataciunas, Deividas
    O'Mahony, Laura
    Zhang, Mike
    Hettiarachchi, Ramith
    Wilson, Joseph
    Machado, Marina
    Moura, Luisa Souza
    Krzeminski, Dominik
    Fadaeil, Hakimeh
    Ergun, Irem
    Okohl, Ifeoma
    Alaagib, Aisha
    Mudannayake, Oshan Ivantha
    Alyafeai, Zaid
    Chien, Vu Minh
    Ruder, Sebastian
    Guthikonda, Surya
    Alghamdil, Emad A.
    Gehrmann, Sebastian
    Muennighoff, Niklas
    Bartolo, Max
    Kreutzer, Julia
    Ustun, Ahmet
    Fadaee, Marzieh
    Hooker, Sara
    PROCEEDINGS OF THE 62ND ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, VOL 1: LONG PAPERS, 2024, : 11521 - 11567
  • [23] Russian Scientific Periodicals in the Directory of Open-Access Journals
    Domnina, T. N.
    SCIENTIFIC AND TECHNICAL INFORMATION PROCESSING, 2018, 45 (04) : 219 - 234
  • [24] Standardized and Open-Access Glaucoma Dataset for Artificial Intelligence Applications
    Steen, Jessica
    Kiefer, Riley
    Ardali, Mahsa
    Abid, Muhammad
    Amjadian, Ehsan
    INVESTIGATIVE OPHTHALMOLOGY & VISUAL SCIENCE, 2023, 64 (08)
  • [25] A Tale of Two Journals: Open-Access and Subscription-Based Publications
    Bourland, J.
    MEDICAL PHYSICS, 2016, 43 (06) : 3693 - 3694
  • [26] Evidence of open access of scientific publications in Google Scholar: A large-scale analysis
    Martin-Martin, Alberto
    Costas, Rodrigo
    van Leeuwen, Thed
    Delgado Lopez-Cozar, Emilio
    JOURNAL OF INFORMETRICS, 2018, 12 (03) : 819 - 841
  • [27] The Datacons Project: An Open-access Dataset of Late Roman Consular Dates
    Dosi, Marco
    JOURNAL OF OPEN HUMANITIES DATA, 2024, 10
  • [28] sPlotOpen - An environmentally balanced, open-access, global dataset of vegetation plots
    Sabatini, Francesco Maria
    Lenoir, Jonathan
    Hattab, Tarek
    Arnst, Elise Aimee
    Chytry, Milan
    Dengler, Juergen
    De Ruffray, Patrice
    Hennekens, Stephan M.
    Jandt, Ute
    Jansen, Florian
    Jimenez-Alfaro, Borja
    Kattge, Jens
    Levesley, Aurora
    Pillar, Valerio D.
    Purschke, Oliver
    Sandel, Brody
    Sultana, Fahmida
    Aavik, Tsipe
    Acic, Svetlana
    Acosta, Alicia T. R.
    Agrillo, Emiliano
    Alvarez, Miguel
    Apostolova, Iva
    Arfin Khan, Mohammed A. S.
    Arroyo, Luzmila
    Attorre, Fabio
    Aubin, Isabelle
    Banerjee, Arindam
    Bauters, Marijn
    Bergeron, Yves
    Bergmeier, Erwin
    Biurrun, Idoia
    Bjorkman, Anne D.
    Bonari, Gianmaria
    Bondareva, Viktoria
    Brunet, Jorg
    Carni, Andraz
    Casella, Laura
    Cayuela, Luis
    Cerny, Tomas
    Chepinoga, Victor
    Csiky, Janos
    Custerevska, Renata
    De Bie, Els
    de Gasper, Andre Luis
    De Sanctis, Michele
    Dimopoulos, Panayotis
    Dolezal, Jiri
    Dziuba, Tetiana
    El-Sheikh, Mohamed Abd El-Rouf Mousa
    GLOBAL ECOLOGY AND BIOGEOGRAPHY, 2021, 30 (09): : 1740 - 1764
  • [29] A reflection on open-access, citation counts, and the future of scientific publishing
    Bosch, Xavier
    ARCHIVUM IMMUNOLOGIAE ET THERAPIAE EXPERIMENTALIS, 2009, 57 (02) : 91 - 93
  • [30] Open Access: Is the Scientific Quality of Biomedical Publications Threatened?
    Barreiro, Esther
    ARCHIVOS DE BRONCONEUMOLOGIA, 2013, 49 (12): : 505 - 506