Excalibur: A new ensemble method based on an optimal combination of aggregation tests for rare-variant association testing for sequencing data

被引:1
作者
Boutry, Simon [1 ,2 ]
Helaers, Raphael [1 ]
Lenaerts, Tom [2 ,3 ,4 ]
Vikkula, Miikka [1 ,5 ]
机构
[1] Univ Louvain, Human Mol Genet, de Duve Inst, Brussels, Belgium
[2] Vrije Univ Brussel, Univ Libre Bruxelles, Interuniv Inst Bioinformat Brussels, Brussels, Belgium
[3] Univ Libre Bruxelles, Machine Learning Grp, Brussels, Belgium
[4] Vrije Univ Brussel, Artificial Intelligence Lab, Brussels, Belgium
[5] WEL Res Inst, WELBIO Dept, Wavre, Belgium
关键词
STATISTICAL TESTS; DISEASE ASSOCIATION; COMMON DISEASES; DETECTING ASSOCIATIONS; GENETIC ASSOCIATION; GENERAL FRAMEWORK; MULTIPLE SNPS; R PACKAGE; POWER; PATHOGENICITY;
D O I
10.1371/journal.pcbi.1011488
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
The development of high-throughput next-generation sequencing technologies and large-scale genetic association studies produced numerous advances in the biostatistics field. Various aggregation tests, i.e. statistical methods that analyze associations of a trait with multiple markers within a genomic region, have produced a variety of novel discoveries. Notwithstanding their usefulness, there is no single test that fits all needs, each suffering from specific drawbacks. Selecting the right aggregation test, while considering an unknown underlying genetic model of the disease, remains an important challenge. Here we propose a new ensemble method, called Excalibur, based on an optimal combination of 36 aggregation tests created after an in-depth study of the limitations of each test and their impact on the quality of result. Our findings demonstrate the ability of our method to control type I error and illustrate that it offers the best average power across all scenarios. The proposed method allows for novel advances in Whole Exome/Genome sequencing association studies, able to handle a wide range of association models, providing researchers with an optimal aggregation analysis for the genetic regions of interest. An increasing number of diseases previously thought to be caused by a mutation in a single gene are now being considered as involving several variants in a small number of genes (i.e. "oligogenic"). There is a limited number of dedicated bioinformatic tools to study such oligogenic causes of diseases. These include so called aggregation tests. Yet, an important challenge is to select the right aggregation test among the various ones that have been developed, as each suffers from different limitations. We have computationally compared 59 aggregation methods to explore their limitations. We found that combining 36 of them results in a more robust method, which we baptized "Excalibur". It can handle a wider range of hypotheses and case-control studies than any of the single methods, while reducing the number of false positive results. Excalibur also provides a comprehensive elucidation of the underlying genetic architecture pertaining to each genomic region under investigation. Thus, it provides a user-friendly, and statistically sound platform to study oligogenic inheritance with the increasing amount of available genetic data.
引用
收藏
页数:26
相关论文
共 102 条
[11]   Analysis of multiple SNPs in a candidate gene or region [J].
Chapman, Juliet ;
Whittaker, John .
GENETIC EPIDEMIOLOGY, 2008, 32 (06) :560-566
[12]   Efficient Variant Set Mixed Model Association Tests for Continuous and Binary Traits in Large-Scale Whole-Genome Sequencing Studies [J].
Chen, Han ;
Huffman, Jennifer E. ;
Brody, Jennifer A. ;
Wang, Chaolong ;
Lee, Seunggeun ;
Li, Zilin ;
Gogarten, Stephanie M. ;
Sofer, Tamar ;
Bielak, Lawrence F. ;
Bis, Joshua C. ;
Blangero, John ;
Bowler, Russell P. ;
Cade, Brian E. ;
Cho, Michael H. ;
Correa, Adolfo ;
Curran, Joanne E. ;
de Vries, Paul S. ;
Glahn, David C. ;
Guo, Xiuqing ;
Johnson, Andrew D. ;
Kardia, Sharon ;
Kooperberg, Charles ;
Lewis, Joshua P. ;
Liu, Xiaoming ;
Mathias, Rasika A. ;
Mitchell, Braxton D. ;
O'Connell, Jeffrey R. ;
Peyser, Patricia A. ;
Post, Wendy S. ;
Reiner, Alex P. ;
Rich, Stephen S. ;
Rotter, Jerome I. ;
Silverman, Edwin K. ;
Smith, Jennifer A. ;
Vasan, Ramachandran S. ;
Wilson, James G. ;
Yanek, Lisa R. ;
Redline, Susan ;
Smith, Nicholas L. ;
Boerwinkle, Eric ;
Borecki, Ingrid B. ;
Cupples, L. Adrienne ;
Laurie, Cathy C. ;
Morrison, Alanna C. ;
Rice, Kenneth M. ;
Lin, Xihong .
AMERICAN JOURNAL OF HUMAN GENETICS, 2019, 104 (02) :260-274
[13]   Control for Population Structure and Relatedness for Binary Traits in Genetic Association Studies via Logistic Mixed Models [J].
Chen, Han ;
Wang, Chaolong ;
Conomos, Matthew P. ;
Stilp, Adrienne M. ;
Li, Zilin ;
Sofer, Tamar ;
Szpiro, Adam A. ;
Chen, Wei ;
Brehm, John M. ;
Celedon, Juan C. ;
Redline, Susan ;
Papanicolaou, George J. ;
Thornton, Timothy A. ;
Laurie, Cathy C. ;
Rice, Kenneth ;
Lin, Xihong .
AMERICAN JOURNAL OF HUMAN GENETICS, 2016, 98 (04) :653-666
[14]   Small Sample Kernel Association Tests for Human Genetic and Microbiome Association Studies [J].
Chen, Jun ;
Chen, Wenan ;
Zhao, Ni ;
Wu, Michael C. ;
Schaid, Daniel J. .
GENETIC EPIDEMIOLOGY, 2016, 40 (01) :5-19
[15]   RVFam: an R package for rare variant association analysis with family data [J].
Chen, Ming-Huei ;
Yang, Qiong .
BIOINFORMATICS, 2016, 32 (04) :624-626
[16]   Recent advances and challenges of rare variant association analysis in the biobank sequencing era [J].
Chen, Wenan ;
Coombes, Brandon J. J. ;
Larson, Nicholas B. B. .
FRONTIERS IN GENETICS, 2022, 13
[17]   A Fast and Noise-Resilient Approach to Detect Rare-Variant Associations With Deep Sequencing Data for Complex Disorders [J].
Cheung, Yee Him ;
Wang, Gao ;
Leal, Suzanne M. ;
Wang, Shuang .
GENETIC EPIDEMIOLOGY, 2012, 36 (07) :675-685
[18]   FARVAT: a family-based rare variant association test [J].
Choi, Sungkyoung ;
Lee, Sungyoung ;
Cichon, Sven ;
Noethen, Markus M. ;
Lange, Christoph ;
Park, Taesung ;
Won, Sungho .
BIOINFORMATICS, 2014, 30 (22) :3197-3205
[19]   Use of unphased multilocus genotype data in indirect association studies [J].
Clayton, D ;
Chapman, J ;
Cooper, J .
GENETIC EPIDEMIOLOGY, 2004, 27 (04) :415-428
[20]   THE COMBINATION OF ESTIMATES FROM DIFFERENT EXPERIMENTS [J].
COCHRAN, WG .
BIOMETRICS, 1954, 10 (01) :101-129