Topic Detection and Multi-word Terms Extraction for Arabic Unvowelized Documents

被引:0
作者
Koulali, Rim [1 ]
Meziane, Ahdelouafi [1 ]
机构
[1] Mohammed 1 Univ, Coll Sci, LARI Lab, Hay Al Quods, Oujda, Morocco
来源
INFORMATION RETRIEVAL TECHNOLOGY | 2011年 / 7097卷
关键词
Topic Detection; Topic Oriented Vocabulary; Mutual Information; Jaccard Indicator; TF-IDF; Multi-Word Terms Extraction; C-value; LLR;
D O I
暂无
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
This paper focuses on Topic Detection (TD) for Arabic Unvowelized documents. Our topic detection system was implemented using two different metrics: adapted TF-IDF and Jaccard indicator. The experiments were conducted while studying the impact of working with stems or roots of words, all the words or nouns only. To enhance the TD system we developed The MWTs extraction prototype to generate MWTs vocabularies. To the best of our knowledge MWTs vocabulary has never been used in arabic documents topic's detection. In this paper we investigate the impact of such use on the quality of topic detection. We used the standard measures: Recall, Precision and F-measure to evaluate the performance of the realized systems on Wattan; an Arabic newspaper corpus.
引用
收藏
页码:614 / 623
页数:10
相关论文
empty
未找到相关数据