A Ranking-Based Text Matching Approach for Plagiarism Detection

被引:3
作者
Kong, Leilei [1 ]
Han, Zhongyuan [1 ]
Qi, Haoliang [2 ]
Lu, Zhimao [3 ]
机构
[1] Heilongjiang Inst Technol, Harbin, Heilongjiang, Peoples R China
[2] State Key Lab Digital Publishing Technol China, Harbin, Heilongjiang, Peoples R China
[3] Dalian Univ Technol, Dalian, Peoples R China
基金
中国国家自然科学基金;
关键词
plagiarism detection; plagiarism text matching; high-obfuscation plagiarism; ranking; meteor; N-GRAMS;
D O I
10.1587/transfun.E101.A.799
中图分类号
TP3 [计算技术、计算机技术];
学科分类号
0812 ;
摘要
This paper addresses the issue of text matching for plagiarism detection. This task aims at identifying the matching plagiarism segments in a pair of suspicious document and its plagiarism source document. All the time, heuristic-based methods are mainly utilized to resolve this problem. But the heuristics rely on the experts' experiences and fail to integrate more features to detect the high obfuscation plagiarism matches. In this paper, a statistical machine learning approach, named the Ranking-based Text Matching Approach for Plagiarism Detection, is proposed to deal with the issues of high obfuscation plagiarism detection. The plagiarism text matching is formalized as a ranking problem, and a pairwise learning to rank algorithm is exploited to identify the most probable plagiarism matches for a given suspicious segment. Especially, the Meteor evaluation metrics of machine translation are subsumed by the proposed method to capture the lexical and semantic text similarity. The proposed method is evaluated on PAN12 and PAN13 text alignment corpus of plagiarism detection and compared to the methods achieved the best performance in PAN12, PAN13 and PAN14. Experimental results demonstrate that the proposed method achieves statistically significantly better performance than the baseline methods in all twelve document collections belonging to five different plagiarism categories. Especially at the PAN12 Artificial-high Obfuscation sub-corpus and PAN13 Summary Obfuscation plagiarism sub-corpus, the main evaluation metrics PlagDet of the proposed method are even 22% and 43% relative improvements than the baselines. Moreover, the efficiency of the proposed method is also better than that of baseline methods.
引用
收藏
页码:799 / 810
页数:12
相关论文
共 33 条
  • [1] Abnar S., 2014, P CLEF 2014 C LABS E
  • [2] Alvi F, 2014, CLEF WORKING NOTES, V1180, P939
  • [3] Understanding Plagiarism Linguistic Patterns, Textual Features, and Detection Methods
    Alzahrani, Salha M.
    Salim, Naomie
    Abraham, Ajith
    [J]. IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS PART C-APPLICATIONS AND REVIEWS, 2012, 42 (02): : 133 - 149
  • [4] [Anonymous], 2014, Proceedings of the Conference and Labs of the Evaluation Forum
  • [5] [Anonymous], 2008, 2008 3rd International Conference on Innovative Computing Information and Control, DOI [10.1109/ICICIC.2008.422, DOI 10.1109/ICICIC.2008.422]
  • [6] [Anonymous], 2002, P ACM SIGKDD KDD 200
  • [7] [Anonymous], 2005, P WORKSH INTR EXTR E
  • [8] Barrón-Cedeño A, 2010, LECT NOTES COMPUT SC, V6008, P687, DOI 10.1007/978-3-642-12116-6_58
  • [9] Barrón-Cedeño A, 2009, LECT NOTES COMPUT SC, V5478, P696, DOI 10.1007/978-3-642-00958-7_69
  • [10] Boosting Algorithm and Meta-Heuristic Based on Genetic Algorithms for Textual Plagiarism Detection
    Bouarara, Hadj Ahmed
    Hamou, Reda Mohamed
    Rahmani, Amine
    Amine, Abdelmalek
    [J]. INTERNATIONAL JOURNAL OF COGNITIVE INFORMATICS AND NATURAL INTELLIGENCE, 2015, 9 (04) : 65 - 87