Outcome-guided Bayesian clustering for disease subtype discovery using high-dimensional transcriptomic data

被引：0

作者：

Meng, Lingsong ^{[1
]}

Huo, Zhiguang ^{[1
,2
]}

机构：

[1] Univ Florida, Dept Biostat, Gainesville, FL USA

[2] 2004 Mowry Rd, Gainesville, FL 32611 USA

来源：

JOURNAL OF APPLIED STATISTICS | 2025年 / 52卷 / 01期

关键词：

Outcome-guided clustering; Bayesian method; Gaussian mixed model; gibbs sampling; INTEGRATED GENOMIC ANALYSIS; SPARSE K-MEANS; BREAST-CANCER; GENE-EXPRESSION; MOLECULAR SUBTYPES; RELEVANT SUBTYPES; MODEL; IDENTIFICATION; SELECTION; SURVIVAL;

D O I：

10.1080/02664763.2024.2362275

中图分类号：

O21 [概率论与数理统计]; C8 [统计学];

学科分类号：

020208 ; 070103 ; 0714 ;

摘要：

Due to the tremendous heterogeneity of disease manifestations, many complex diseases that were once thought to be single diseases are now considered to have disease subtypes. Disease subtyping analysis, that is the identification of subgroups of patients with similar characteristics, is the first step to accomplish precision medicine. With the advancement of high-throughput technologies, omics data offers unprecedented opportunity to reveal disease subtypes. As a result, unsupervised clustering analysis has been widely used for this purpose. Though promising, the subtypes obtained from traditional quantitative approaches may not always be clinically meaningful (i.e. correlate with clinical outcomes). On the other hand, the collection of rich clinical data in modern epidemiology studies has the great potential to facilitate the disease subtyping process via omics data and to discovery clinically meaningful disease subtypes. Thus, we developed an outcome-guided Bayesian clustering (GuidedBayesianClustering) method to fully integrate the clinical data and the high-dimensional omics data. A Gaussian mixed model framework was applied to perform sample clustering; a spike-and-slab prior was utilized to perform gene selection; a mixture model prior was employed to incorporate the guidance from a clinical outcome variable; and a decision framework was adopted to infer the false discovery rate of the selected genes. We deployed conjugate priors to facilitate efficient Gibbs sampling. Our proposed full Bayesian method is capable of simultaneously (i) obtaining sample clustering (disease subtype discovery); (ii) performing feature selection (select genes related to the disease subtype); and (iii) utilizing clinical outcome variable to guide the disease subtype discovery. The superior performance of the GuidedBayesianClustering was demonstrated through simulations and applications of breast cancer expression data and Alzheimer's disease. An R package has been made publicly available on GitHub to improve the applicability of our method.

引用

页码：183 / 207

页数：25

共 50 条

[31] ForestSubtype: a cancer subtype identifying approach based on high-dimensional genomic data and a parallel random forest
Luo, Junwei
Feng, Yading
Wu, Xuyang
Li, Ruimin
Shi, Jiawei
Chang, Wenjing
Wang, Junfeng
BMC BIOINFORMATICS, 2023, 24 (01)
[32] Model-based regression clustering for high-dimensional data: application to functional data
Devijver, Emilie
ADVANCES IN DATA ANALYSIS AND CLASSIFICATION, 2017, 11 (02) : 243 - 279
[33] Clustering electricity consumers using high-dimensional regression mixture models
Devijver, Emilie
Goude, Yannig
Poggi, Jean-Michel
APPLIED STOCHASTIC MODELS IN BUSINESS AND INDUSTRY, 2020, 36 (01) : 159 - 177
[34] Scalable spatio-temporal Bayesian analysis of high-dimensional electroencephalography data
Mohammed, Shariq
Dey, Dipak K.
CANADIAN JOURNAL OF STATISTICS-REVUE CANADIENNE DE STATISTIQUE, 2021, 49 (01): : 107 - 128
[35] Penalized mixtures of factor analyzers with application to clustering high-dimensional microarray data
Xie, Benhuai
Pan, Wei
Shen, Xiaotong
BIOINFORMATICS, 2010, 26 (04) : 501 - 508
[36] Robust clustering of noisy high-dimensional gene expression data for patients subtyping
Coretto, Pietro
Serra, Angela
Tagliaferri, Roberto
BIOINFORMATICS, 2018, 34 (23) : 4064 - 4072
[37] Analyzing high-dimensional cytometry data using FlowSOM
Quintelier, Katrien
Couckuyt, Artuur
Emmaneel, Annelies
Aerts, Joachim
Saeys, Yvan
Van Gassen, Sofie
NATURE PROTOCOLS, 2021, 16 (08) : 3775 - 3801
[38] Bayesian penalized cumulative logit model for high-dimensional data with an ordinal response
Zhang, Yiran
Archer, Kellie J.
STATISTICS IN MEDICINE, 2021, 40 (06) : 1453 - 1481
[39] Severe testing with high-dimensional omics data for enhancing biomedical scientific discovery
Emmert-Streib, Frank
NPJ SYSTEMS BIOLOGY AND APPLICATIONS, 2022, 8 (01)
[40] Genomic data analysis using a two stage expectation propagation algorithm for analysis of sparse Bayesian high-dimensional instrumental variables regression
Amini, Morteza
COMMUNICATIONS IN STATISTICS-SIMULATION AND COMPUTATION, 2024, 53 (05) : 2351 - 2365

← 1 2 3 4 5 →