Unsupervised Feature Selection for Latent Dirichlet Allocation

被引:0
作者
Xu Weiran [1 ]
Du Gang [1 ]
Chen Guang [1 ]
Guo Jun [1 ]
Yang Jie [1 ]
机构
[1] Beijing Univ Posts & Telecommun, Sch Informat & Commun Engn, Beijing 100876, Peoples R China
关键词
pattern recognition; unsupervised feature selection; Latent Dirichlet Allocation; general topic; special topic;
D O I
暂无
中图分类号
TN [电子技术、通信技术];
学科分类号
0809 ;
摘要
As a generative model, Latent Dirichlet Allocation Model, which lacks optimization of topics' discrimination capability focuses on how to generate data, This paper aims to improve the discrimination capability through unsupervised feature selection. Theoretical analysis shows that the discrimination capability of a topic is limited by the discrimination capability of its representative words. The discrimination capability of a word is approximated by the Information Gain of the word for topics, which is used to distinguish between "general word" and special word" in LDA topics. Therefore, we add a constraint to the LDA objective function to let the "general words" only happen in "general topics" other than "special topics". Then a heuristic algorithm is presented to get the solution. Experiments show that this method can not only improve the information gain of topics, but also make the topics easier to understand by human.
引用
收藏
页码:54 / 62
页数:9
相关论文
共 17 条
[1]  
[Anonymous], 1999, CLAIMING PLACE P 15
[2]   Latent Dirichlet allocation [J].
Blei, DM ;
Ng, AY ;
Jordan, MI .
JOURNAL OF MACHINE LEARNING RESEARCH, 2003, 3 (4-5) :993-1022
[3]  
BOUTSIDIS C, 2008, P KNOWL DISC DAT MIN
[4]  
DANIEL DL, 2000, P C ADV NEUR INF PRO, V13, P556
[5]  
DEERWESTER S, 1990, J AM SOC INFORM SCI, V41, P391, DOI 10.1002/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO
[6]  
2-9
[7]  
DING CHRIS, 2008, COMPUTATIONAL STAT D, V8, P3913
[8]  
Dy JG, 2004, J MACH LEARN RES, V5, P845
[9]  
Guyon I., 2003, J MACH LEARN RES, V3, P1157
[10]  
HONG YI, 2008, PATTERN RECOGN, V9, P2742