Distributed Bayesian Matrix Decomposition for Big Data Mining and Clustering

被引：7

作者：

Zhang, Chihao ^{[1
,2
,3
]}

Yang, Yang ^{[4
]}

Zhou, Wei ^{[4
]}

Zhang, Shihua ^{[1
,2
,3
]}

机构：

[1] Chinese Acad Sci, Acad Math & Syst Sci, RCSDS, NCMIS,CEMS, Beijing 100190, Peoples R China

[2] Univ Chinese Acad Sci, Sch Math Sci, Beijing 100049, Peoples R China

[3] Chinese Acad Sci, Ctr Excellence Anim Evolut & Genet, Kunming 650223, Yunnan, Peoples R China

[4] Yunnan Univ, Sch Software, Kunming 650504, Yunnan, Peoples R China

来源：

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING | 2022年 / 34卷 / 08期

基金：

中国国家自然科学基金;

关键词：

Matrix decomposition; Bayes methods; Big Data; Principal component analysis; Distributed databases; Data mining; Clustering algorithms; Distributed algorithm; bayesian matrix decomposition; clustering; big data; data mining; FACTORIZATION; MODEL;

D O I：

10.1109/TKDE.2020.3029582

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Matrix decomposition is one of the fundamental tools to discover knowledge from big data generated by modern applications. However, it is still inefficient or infeasible to process very big data using such a method in a single machine. Moreover, big data are often distributedly collected and stored on different machines. Thus, such data generally bear strong heterogeneous noise. It is essential and useful to develop distributed matrix decomposition for big data analytics. Such a method should scale up well, model the heterogeneous noise, and address the communication issue in a distributed system. To this end, we propose a distributed Bayesian matrix decomposition model (DBMD) for big data mining and clustering. Specifically, we adopt three strategies to implement the distributed computing including 1) the accelerated gradient descent, 2) the alternating direction method of multipliers (ADMM), and 3) the statistical inference. We investigate the theoretical convergence behaviors of these algorithms. To address the heterogeneity of the noise, we propose an optimal plug-in weighted average that reduces the variance of the estimation. Synthetic experiments validate our theoretical results, and real-world experiments show that our algorithms scale up well to big data and achieves superior or competing performance compared to two typical distributed methods including Scalable-NMF and scalable k-means++.

引用

页码：3701 / 3713

页数：13

共 50 条

[1] Large-Scale Distributed Bayesian Matrix Factorization using Stochastic Gradient MCMC
Ahn, Sungjin
Korattikara, Anoop
Liu, Nathan
Rajan, Suju
Welling, Max
[J]. KDD'15: PROCEEDINGS OF THE 21ST ACM SIGKDD INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA MINING, 2015, : 9 - 18
[2] [Anonymous], 2014, P INT C NEUR INF PRO
[3] [Anonymous], 2011, Acm T. Intel. Syst. Tec., DOI DOI 10.1145/1961189.1961199
[4] Arthur D, 2007, PROCEEDINGS OF THE EIGHTEENTH ANNUAL ACM-SIAM SYMPOSIUM ON DISCRETE ALGORITHMS, P1027
[5] Scalable K-Means++
Bahmani, Bahman
Moseley, Benjamin
Vattani, Andrea
Kumar, Ravi
Vassilvitskii, Sergei
[J]. PROCEEDINGS OF THE VLDB ENDOWMENT, 2012, 5 (07): : 622 - 633
[6] A Fast Iterative Shrinkage-Thresholding Algorithm for Linear Inverse Problems
Beck, Amir
Teboulle, Marc
[J]. SIAM JOURNAL ON IMAGING SCIENCES, 2009, 2 (01): : 183 - 202
[7] Bishop C. M., 2006, PATTERN RECOGN
[8] Bishop CM, 1999, ADV NEUR IN, V11, P382
[9] Distributed optimization and statistical learning via the alternating direction method of multipliers
Boyd S.
Parikh N.
Chu E.
Peleato B.
Eckstein J.
[J]. Foundations and Trends in Machine Learning, 2010, 3 (01): : 1 - 122
[10] Graph Regularized Nonnegative Matrix Factorization for Data Representation
Cai, Deng
He, Xiaofei
Han, Jiawei
Huang, Thomas S.
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2011, 33 (08) : 1548 - 1560

← 1 2 3 4 5 →