Revisiting K-Means and Topic Modeling, a Comparison Study to Cluster Arabic Documents

被引:55
作者
Alhawarat, M. [1 ]
Hegazi, M. [1 ]
机构
[1] Prince Sattam Bin Abdulaziz Univ, Dept Comp Sci, Al Kharj 11942, Saudi Arabia
来源
IEEE ACCESS | 2018年 / 6卷
关键词
Clustering text documents; K-means; Arabic language; topic modeling; latent Dirichlet allocation (LDA); TEXTS;
D O I
10.1109/ACCESS.2018.2852648
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Clustering Arabic text documents is of high importance for many natural language technologies. This paper uses a combined method to cluster Arabic text documents. Mainly, we use generative models and clustering techniques. The study uses latent Dirichlet allocation and k-means clustering algorithm and applies them to a news data set used in previous similar studies. The aim of this paper is twofold: it first shows that normalizing the weights in the vector space, for the document-term matrix of the text documents, dramatically improves the quality of clusters and hence the accuracy of clustering when using k-means algorithm. The results are compared to a recent study on clustering Arabic text documents. Second, it shows that the combined method is superior in terms of clustering quality for Arabic text documents according to external measures, such as purity, F-measure, entropy, accuracy, and other measures. It is shown in this paper that the purity of the combined method is 0.933 compared to 0.82 for k-means algorithm, and these figures are higher in comparison to a recent similar study. This is also confirmed by the other used validation measures. The correctness of the combined method is then confirmed using different Arabic data sets.
引用
收藏
页码:42740 / 42749
页数:10
相关论文
共 61 条
  • [1] Abuaiadah D., 2014, INT J COMPUTER APPL, V101, P31, DOI [10.5120/17701-8680, DOI 10.5120/17701-8680]
  • [3] Agrawal R., 1998, SIGMOD Record, V27, P94, DOI 10.1145/276305.276314
  • [4] Al-Sarrayrih H. S., 2009, P INT C INF TECHN IC
  • [5] Correcting Jaccard and other similarity indices for chance agreement in cluster analysis
    Albatineh, Ahmed N.
    Niewiadomska-Bugaj, Magdalena
    [J]. ADVANCES IN DATA ANALYSIS AND CLASSIFICATION, 2011, 5 (03) : 179 - 200
  • [6] Alhawarat M, 2015, INT J ADV COMPUT SC, V6, P288
  • [7] A comparison of extrinsic clustering evaluation metrics based on formal constraints
    Amigo, Enrique
    Gonzalo, Julio
    Artiles, Javier
    Verdejo, Felisa
    [J]. INFORMATION RETRIEVAL, 2009, 12 (04): : 461 - 486
  • [8] [Anonymous], ADV COMPUTATIONAL SC
  • [9] [Anonymous], 2008, Introduction to information retrieval
  • [10] [Anonymous], 2007, INT J INF TECHNOL