Cluster and Feature Modeling from Combinatorial Stochastic Processes

被引:18
作者
Broderick, Tamara [1 ]
Jordan, Michael I. [1 ,2 ]
Pitman, Jim [1 ,3 ]
机构
[1] Univ Calif Berkeley, Dept Stat, Berkeley, CA 94720 USA
[2] Univ Calif Berkeley, Dept EECS, Berkeley, CA 94720 USA
[3] Univ Calif Berkeley, Dept Math, Berkeley, CA 94720 USA
基金
美国国家科学基金会;
关键词
Cluster; feature; Dirichlet process; beta process; Chinese restaurant process; Indian buffet process; nonparametric; Bayesian; combinatorial stochastic process; CHAIN MONTE-CARLO; STICK-BREAKING; SAMPLING METHODS; BETA PROCESSES; DIRICHLET; DISTRIBUTIONS; ESTIMATORS; INFERENCE;
D O I
10.1214/13-STS434
中图分类号
O21 [概率论与数理统计]; C8 [统计学];
学科分类号
020208 ; 070103 ; 0714 ;
摘要
One of the focal points of the modern literature on Bayesian nonparametrics has been the problem of clustering, or partitioning, where each data point is modeled as being associated with one and only one of some collection of groups called clusters or partition blocks. Underlying these Bayesian nonparametric models are a set of interrelated stochastic processes, most notably the Dirichlet process and the Chinese restaurant process. In this paper we provide a formal development of an analogous problem, called feature modeling, for associating data points with arbitrary nonnegative integer numbers of groups, now called features or topics. We review the existing combinatorial stochastic process representations for the clustering problem and develop analogous representations for the feature modeling problem. These representations include the beta process and the Indian buffet process as well as new representations that provide insight into the connections between these processes. We thereby bring the same level of completeness to the treatment of Bayesian nonparametric feature modeling that has previously been achieved for Bayesian nonparametric clustering.
引用
收藏
页码:289 / 312
页数:24
相关论文
共 55 条
  • [1] [Anonymous], TECHNICAL REPORT
  • [2] [Anonymous], DIFFUSIONS MARKOV PR
  • [3] [Anonymous], INT C MACH LEARN HAI
  • [4] [Anonymous], 520 U CAL
  • [5] [Anonymous], P INT C ART INT STAT
  • [6] [Anonymous], 1985, Ecole d'Ete de Probabilites de Saint-Flour XIII
  • [7] [Anonymous], BAYESIAN AN IN PRESS
  • [8] [Anonymous], SUBORDINATORS UNPUB
  • [9] [Anonymous], P INT C ART INT STAT
  • [10] [Anonymous], ARXIV11111802