Due to the tremendous heterogeneity of disease manifestations, many complex diseases that were once thought to be single diseases are now considered to have disease subtypes. Disease subtyping analysis, that is the identification of subgroups of patients with similar characteristics, is the first step to accomplish precision medicine. With the advancement of high-throughput technologies, omics data offers unprecedented opportunity to reveal disease subtypes. As a result, unsupervised clustering analysis has been widely used for this purpose. Though promising, the subtypes obtained from traditional quantitative approaches may not always be clinically meaningful (i.e. correlate with clinical outcomes). On the other hand, the collection of rich clinical data in modern epidemiology studies has the great potential to facilitate the disease subtyping process via omics data and to discovery clinically meaningful disease subtypes. Thus, we developed an outcome-guided Bayesian clustering (GuidedBayesianClustering) method to fully integrate the clinical data and the high-dimensional omics data. A Gaussian mixed model framework was applied to perform sample clustering; a spike-and-slab prior was utilized to perform gene selection; a mixture model prior was employed to incorporate the guidance from a clinical outcome variable; and a decision framework was adopted to infer the false discovery rate of the selected genes. We deployed conjugate priors to facilitate efficient Gibbs sampling. Our proposed full Bayesian method is capable of simultaneously (i) obtaining sample clustering (disease subtype discovery); (ii) performing feature selection (select genes related to the disease subtype); and (iii) utilizing clinical outcome variable to guide the disease subtype discovery. The superior performance of the GuidedBayesianClustering was demonstrated through simulations and applications of breast cancer expression data and Alzheimer's disease. An R package has been made publicly available on GitHub to improve the applicability of our method.
机构:
Seoul Natl Univ, Res Inst Agr & Life Sci, Seoul, South Korea
Seoul Natl Univ, Dept Agr Biotechnol, Seoul, South KoreaSeoul Natl Univ, Res Inst Agr & Life Sci, Seoul, South Korea
Kim, Soo-Jin
Ha, Jung-Woo
论文数: 0引用数: 0
h-index: 0
机构:
NAVER Corp, Clova AI Res, Seongnam, South KoreaSeoul Natl Univ, Res Inst Agr & Life Sci, Seoul, South Korea
Ha, Jung-Woo
Kim, Heebal
论文数: 0引用数: 0
h-index: 0
机构:
Seoul Natl Univ, Res Inst Agr & Life Sci, Seoul, South Korea
Seoul Natl Univ, Dept Agr Biotechnol, Seoul, South Korea
C&K Genom, Seoul, South KoreaSeoul Natl Univ, Res Inst Agr & Life Sci, Seoul, South Korea
Kim, Heebal
Zhang, Byoung-Tak
论文数: 0引用数: 0
h-index: 0
机构:
Seoul Natl Univ, Sch Comp Sci & Engn, Seoul, South KoreaSeoul Natl Univ, Res Inst Agr & Life Sci, Seoul, South Korea
机构:
Univ Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USAUniv Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USA
Tamaresis, John S.
Irwin, Juan C.
论文数: 0引用数: 0
h-index: 0
机构:
Univ Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USAUniv Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USA
Irwin, Juan C.
Goldfien, Gabriel A.
论文数: 0引用数: 0
h-index: 0
机构:
Univ Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USAUniv Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USA
Goldfien, Gabriel A.
Rabban, Joseph T.
论文数: 0引用数: 0
h-index: 0
机构:
Univ Calif San Francisco, Dept Pathol, San Francisco, CA 94143 USAUniv Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USA
Rabban, Joseph T.
Burney, Richard O.
论文数: 0引用数: 0
h-index: 0
机构:
Madigan Healthcare Syst, Dept Obstet & Gynecol & Clin Invest, Tacoma, WA 98431 USAUniv Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USA
Burney, Richard O.
Nezhat, Camran
论文数: 0引用数: 0
h-index: 0
机构:
Stanford Univ, Dept Obstet & Gynecol, Stanford, CA 94024 USAUniv Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USA
Nezhat, Camran
DePaolo, Louis V.
论文数: 0引用数: 0
h-index: 0
机构:
NIH, Fertil & Infertil Branch, Eunice Kennedy Shriver Natl Inst Child Hlth & Hum, Bethesda, MD 20892 USAUniv Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USA
DePaolo, Louis V.
Giudice, Linda C.
论文数: 0引用数: 0
h-index: 0
机构:
Univ Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USAUniv Calif San Francisco, Dept Obstet Gynecol & Reprod Sci, Ctr Reprod Sci, San Francisco, CA 94143 USA