A novel unsupervised corpus-based stemming technique using lexicon and corpus statistics

被引:19
作者
Singh, Jasmeet [1 ]
Gupta, Vishal [2 ]
机构
[1] Thapar Inst Engn & Technol, Patiala, Punjab, India
[2] Panjab Univ, Univ Inst Engn & Technol, Chandigarh, India
关键词
Stemming; Inflection; Morphology; Corpus; Information retrieval; Natural language processing; TEXT;
D O I
10.1016/j.knosys.2019.05.025
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Word Stemming is a widely used mechanism in the fields of Natural Language Processing, Information Retrieval, and Language Modeling. Language-independent stemmers discover classes of morphologically related words from the ambient corpus without using any language related rules. In this article, we proposed a fully unsupervised language-independent text stemming technique that clusters morphologically related words from the corpus of the language using both lexical and co-occurrence features such as lexical similarity, suffix knowledge, and co-occurrence similarity. The method applies to a wide range of inflectional languages as it identifies morphological variants formed through different linguistic processes such as affixation, compounding, conversion, etc. The proposed approach has been tested in Information Retrieval application for four languages (English, Marathi, Hungarian, and Bengali) using standard TREC, CLEF, and FIRE test collections. A significant improvement over word-based retrieval, five other corpus-based stemmers, and rule-based stemmers has been achieved in all the languages. Besides, information retrieval, the proposed approach has also been tested in text classification and inflection removal tasks. Our algorithm excelled over other baseline methods in all the test scenarios. Thus, we successfully achieved the objective of developing a multipurpose stemming algorithm that cannot only be used for information retrieval task but also for non-traditional tasks such as text classification, sentiment analysis, inflection removal, etc. (C) 2019 Elsevier B.V. All rights reserved.
引用
收藏
页码:147 / 162
页数:16
相关论文
共 57 条
[1]   Probabilistic models of information retrieval based on measuring the divergence from randomness [J].
Amati, G ;
Van Rijsbergen, CJ .
ACM TRANSACTIONS ON INFORMATION SYSTEMS, 2002, 20 (04) :357-389
[2]  
[Anonymous], 1994, WETENSCHAPPELIJKE BI
[3]  
[Anonymous], COGN COMPUT
[4]  
[Anonymous], NODALIDA 2003
[5]  
Arif S. M., 2017, P INT C COMP MATH ST, P93
[6]   A probabilistic model for stemmer generation [J].
Bacchin, M ;
Ferro, N ;
Melucci, M .
INFORMATION PROCESSING & MANAGEMENT, 2005, 41 (01) :121-137
[7]  
Baker L. D., 1998, Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, P96, DOI 10.1145/290941.290970
[8]  
Baroni Marco., 2002, Proceedings of the ACL-02 Workshop on Morphological and Phonological Learning, V6, P48, DOI [10.3115/1118647.1118653, DOI 10.48550/ARXIV.CS/0205006, https://doi.org/10.48550/arXiv.cs/0205006]
[9]   Stemming via distribution-based word segregation for classification and retrieval [J].
Bhamidipati, Narayan L. ;
Pal, Sankar K. .
IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS PART B-CYBERNETICS, 2007, 37 (02) :350-360
[10]  
Biba M., 2014, Recent advances in intelligent informatics, P185, DOI DOI 10.1007/978-3-319-01778-5_19