Selecting Priors for Latent Dirichlet Allocation

被引:17
作者
Syed, Shaheen [1 ]
Spruit, Marco [1 ]
机构
[1] Univ Utrecht, Dept Informat & Comp Sci, Utrecht, Netherlands
来源
2018 IEEE 12TH INTERNATIONAL CONFERENCE ON SEMANTIC COMPUTING (ICSC) | 2018年
基金
欧盟地平线“2020”;
关键词
D O I
10.1109/ICSC.2018.00035
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Latent Dirichlet Allocation (LDA) has gained much attention from researchers and is increasingly being applied to uncover underlying semantic structures from a variety of corpora. However, nearly all researchers use symmetrical Dirichlet priors, often unaware of the underlying practical implications that they bear. This research is the first to explore symmetrical and asymmetrical Dirichlet priors on topic coherence and human topic ranking when uncovering latent semantic structures from scientific research articles. More specifically, we examine the practical effects of several classes of Dirichlet priors on 2000 LDA models created from abstract and full-text research articles. Our results show that symmetrical or asymmetrical priors on the document-topic distribution or the topic-word distribution for full-text data have little effect on topic coherence scores and human topic ranking. In contrast, asymmetrical priors on the document-topic distribution for abstract data show a significant increase in topic coherence scores and improved human topic ranking compared to a symmetrical prior. Symmetrical or asymmetrical priors on the topic-word distribution show no real benefits for both abstract and full-text data.
引用
收藏
页码:194 / 202
页数:9
相关论文
共 38 条
[11]   Mixed-membership models of scientific publications [J].
Erosheva, E ;
Fienberg, S ;
Lafferty, J .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 2004, 101 :5220-5227
[12]   Latent Semantic Analysis: five methodological recommendations [J].
Evangelopoulos, Nicholas ;
Zhang, Xiaoni ;
Prybutok, Victor R. .
EUROPEAN JOURNAL OF INFORMATION SYSTEMS, 2012, 21 (01) :70-86
[13]  
Fergus R, 2005, IEEE I CONF COMP VIS, P1816
[14]  
Hall D., 2008, P 2008 C EMP METH NA, P363, DOI DOI 10.3115/1613715.1613763
[15]   DISTRIBUTIONAL STRUCTURE [J].
Harris, Zellig S. .
WORD-JOURNAL OF THE INTERNATIONAL LINGUISTIC ASSOCIATION, 1954, 10 (2-3) :146-162
[16]  
Heinrich Gregor, 2005, TECHNICAL REPORT
[17]  
Hoffman M. D., 2010, P ADV NEUR INF PROC, P856
[18]   Probabilistic latent semantic indexing [J].
Hofmann, T .
SIGIR'99: PROCEEDINGS OF 22ND INTERNATIONAL CONFERENCE ON RESEARCH AND DEVELOPMENT IN INFORMATION RETRIEVAL, 1999, :50-57
[19]  
Huang J., 2005, Tech. rep., CMU Technique Report
[20]  
Kim Samuel, 2009, 2009 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), P37, DOI 10.1109/ASPAA.2009.5346483