Detecting differentially expressed genes in microarrays using Bayesian model selection

被引:108
作者
Ishwaran, H [1 ]
Rao, JS
机构
[1] Cleveland Clin Fdn, Dept Biostat & Epidemiol Wb4, Cleveland, OH 44195 USA
[2] Case Western Reserve Univ, Dept Epidemiol & Biostat, Cleveland, OH 44106 USA
关键词
Bayesian analysis of variance for microarrays; false discovery rate; false nondiscovery rate; heteroscedasticity; ridge; regression; Shrinkage; variance stabilizing transform; weighted regression;
D O I
10.1198/016214503000224
中图分类号
O21 [概率论与数理统计]; C8 [统计学];
学科分类号
020208 ; 070103 ; 0714 ;
摘要
DNA microarrays open up a broad new horizon for investigators interested in studying the genetic determinants of disease. The high throughput nature of these arrays, where differential expression for thousands of genes can be measured simultaneously, creates an enormous wealth of information, but also poses a challenge for data analysis because of the large multiple testing problem involved. The solution has generally been to focus on optimizing false-discovery rates while sacrificing power. The drawback of this approach is that more subtle expression differences will be missed that might give investigators more insight into the genetic environment necessary for a disease process to take hold. We introduce a new method for detecting differentially expressed genes based on a high-dimensional model selection technique, Bayesian ANOVA for microarrays (BAM), which strikes a balance between false rejections and false nonrejections. The basis of the new approach involves a weighted average of generalized ridge regression estimates that provides the benefits of using shrinkage estimation combined with model averaging. A simple graphical tool based on the amount of shrinkage is developed to visualize the trade-off between low false-discovery rates and finding more genes. Simulations are used to illustrate BAM's performance, and the method is applied to a large database of colon cancer gene expression data. Our working hypothesis in the colon cancer analysis is that large differential expressions may not be the only ones contributing to metastasis-in fact, moderate changes in expression of genes may be involved in modifying the genetic environment to a sufficient extent for metastasis to occur. A functional biological analysis of gene effects found by BAM, but not other false-discovery-based approaches, lends support to this hypothesis.
引用
收藏
页码:438 / 455
页数:18
相关论文
共 31 条
[1]  
BARBIERI MM, 2002, 0202 ISDS
[2]  
Benjamini Y, 2001, ANN STAT, V29, P1165
[3]   CONTROLLING THE FALSE DISCOVERY RATE - A PRACTICAL AND POWERFUL APPROACH TO MULTIPLE TESTING [J].
BENJAMINI, Y ;
HOCHBERG, Y .
JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY, 1995, 57 (01) :289-300
[4]   Exploring the new world of the genome with DNA microarrays [J].
Brown, PO ;
Botstein, D .
NATURE GENETICS, 1999, 21 (Suppl 1) :33-37
[5]   Microarray expression profiling identifies genes with altered expression in HDL-deficient mice [J].
Callow, MJ ;
Dudoit, S ;
Gong, EL ;
Speed, TP ;
Rubin, EM .
GENOME RESEARCH, 2000, 10 (12) :2022-2029
[6]  
Chow Y., 1978, PROBABILITY THEORY I
[7]  
Cohen Alfred M., 1997, P1144
[8]  
Durbin B P, 2002, Bioinformatics, V18 Suppl 1, pS105
[9]   Empirical Bayes analysis of a microarray experiment [J].
Efron, B ;
Tibshirani, R ;
Storey, JD ;
Tusher, V .
JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2001, 96 (456) :1151-1160
[10]   Analysis of expressed sequence tags indicates 35,000 human genes [J].
Ewing, B ;
Green, P .
NATURE GENETICS, 2000, 25 (02) :232-234