A Persian Part-Of-Speech Tagger Based on Morphological Analysis

被引:0
作者
Mohseni, Mahdi [1 ]
Minaei-bidgoli, Behrouz [1 ]
机构
[1] Iran Univ Sci & Technol, Tehran, Iran
来源
LREC 2010 - SEVENTH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION | 2010年
关键词
D O I
暂无
中图分类号
H [语言、文字];
学科分类号
05 ;
摘要
This paper describes a method based on morphological analysis of words for a Persian Part-Of-Speech (POS) tagging system. This is a main part of a process for expanding a large Persian corpus called Peyekare (or Textual Corpus of Persian Language). Peykare is arranged into two parts: annotated and unannotated parts. We use the annotated part in order to create an automatic morphological analyzer, a main segment of the system. Morphosyntactic features of Persian words cause two problems: the number of tags is increased in the corpus (586 tags) and the form of the words is changed. This high number of tags debilitates any taggers to work efficiently. From other side the change of word forms reduces the frequency of words with the same lemma; and the number of words belonging to a specific tag reduces as well. This problem also has a bad effect on statistical taggers. The morphological analyzer by removing the problems helps the tagger to cover a large number of tags in the corpus. Using a Markov tagger the method is evaluated on the corpus. The experiments show the efficiency of the method in Persian POS tagging.
引用
收藏
页码:1253 / 1257
页数:5
相关论文
共 19 条
[1]  
[Anonymous], 1996, P 4 WORKSHOP VERY LA
[2]  
Assi S. M., 2000, INT J CORPUS LINGUIS, V5, P69
[3]  
Assi S. M., 1997, INT J LEXICOGR, V10, P5
[4]  
Bijankhan M., 2002, PERSIAN LANGUAGE MOD
[5]  
CASTOR A, 1992, APPL INTELL, V2, P37
[6]  
CHARNIAK E, 1993, PROCEEDINGS OF THE ELEVENTH NATIONAL CONFERENCE ON ARTIFICIAL INTELLIGENCE, P784
[7]  
Chercheur J. L., 1994, CASE BASED REASONING
[8]  
Grandchercheur L.B., 1983, FONDEMENT SCI COGNIT, P6
[9]  
Kupiec J., 1992, Computer Speech and Language, V6, P225, DOI 10.1016/0885-2308(92)90019-Z
[10]  
Leech G., 1999, Syntactic word-class tagging, P55