Excalibur: A new ensemble method based on an optimal combination of aggregation tests for rare-variant association testing for sequencing data

被引:1
作者
Boutry, Simon [1 ,2 ]
Helaers, Raphael [1 ]
Lenaerts, Tom [2 ,3 ,4 ]
Vikkula, Miikka [1 ,5 ]
机构
[1] Univ Louvain, Human Mol Genet, de Duve Inst, Brussels, Belgium
[2] Vrije Univ Brussel, Univ Libre Bruxelles, Interuniv Inst Bioinformat Brussels, Brussels, Belgium
[3] Univ Libre Bruxelles, Machine Learning Grp, Brussels, Belgium
[4] Vrije Univ Brussel, Artificial Intelligence Lab, Brussels, Belgium
[5] WEL Res Inst, WELBIO Dept, Wavre, Belgium
关键词
STATISTICAL TESTS; DISEASE ASSOCIATION; COMMON DISEASES; DETECTING ASSOCIATIONS; GENETIC ASSOCIATION; GENERAL FRAMEWORK; MULTIPLE SNPS; R PACKAGE; POWER; PATHOGENICITY;
D O I
10.1371/journal.pcbi.1011488
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
The development of high-throughput next-generation sequencing technologies and large-scale genetic association studies produced numerous advances in the biostatistics field. Various aggregation tests, i.e. statistical methods that analyze associations of a trait with multiple markers within a genomic region, have produced a variety of novel discoveries. Notwithstanding their usefulness, there is no single test that fits all needs, each suffering from specific drawbacks. Selecting the right aggregation test, while considering an unknown underlying genetic model of the disease, remains an important challenge. Here we propose a new ensemble method, called Excalibur, based on an optimal combination of 36 aggregation tests created after an in-depth study of the limitations of each test and their impact on the quality of result. Our findings demonstrate the ability of our method to control type I error and illustrate that it offers the best average power across all scenarios. The proposed method allows for novel advances in Whole Exome/Genome sequencing association studies, able to handle a wide range of association models, providing researchers with an optimal aggregation analysis for the genetic regions of interest. An increasing number of diseases previously thought to be caused by a mutation in a single gene are now being considered as involving several variants in a small number of genes (i.e. "oligogenic"). There is a limited number of dedicated bioinformatic tools to study such oligogenic causes of diseases. These include so called aggregation tests. Yet, an important challenge is to select the right aggregation test among the various ones that have been developed, as each suffers from different limitations. We have computationally compared 59 aggregation methods to explore their limitations. We found that combining 36 of them results in a more robust method, which we baptized "Excalibur". It can handle a wider range of hypotheses and case-control studies than any of the single methods, while reducing the number of false positive results. Excalibur also provides a comprehensive elucidation of the underlying genetic architecture pertaining to each genomic region under investigation. Thus, it provides a user-friendly, and statistically sound platform to study oligogenic inheritance with the increasing amount of available genetic data.
引用
收藏
页数:26
相关论文
共 102 条
[1]  
Adzhubei Ivan, 2013, Curr Protoc Hum Genet, VChapter 7, DOI 10.1002/0471142905.hg0720s76
[2]   TESTS FOR LINEAR TRENDS IN PROPORTIONS AND FREQUENCIES [J].
ARMITAGE, P .
BIOMETRICS, 1955, 11 (03) :375-386
[3]   Rare Variant Association Analysis Methods for Complex Traits [J].
Asimit, Jennifer ;
Zeggini, Eleftheria .
ANNUAL REVIEW OF GENETICS, VOL 44, 2010, 44 :293-308
[4]   ARIEL and AMELIA: Testing for an Accumulation of Rare Variants Using Next-Generation Sequencing Data [J].
Asimit, Jennifer L. ;
Day-Williams, Aaron G. ;
Morris, Andrew P. ;
Zeggini, Eleftheria .
HUMAN HEREDITY, 2012, 73 (02) :84-94
[5]  
Banerjee Amitav, 2009, Ind Psychiatry J, V18, P127, DOI 10.4103/0972-6748.62274
[6]   Comparison of Statistical Tests for Disease Association With Rare Variants [J].
Basu, Saonli ;
Pan, Wei .
GENETIC EPIDEMIOLOGY, 2011, 35 (07) :606-619
[7]   FREGAT: an R package for region-based association analysis [J].
Belonogova, Nadezhda M. ;
Svishcheva, Gulnara R. ;
Axenovich, Tatiana I. .
BIOINFORMATICS, 2016, 32 (15) :2392-2393
[8]  
Berstein Y., 2018, Detection of rare disease-related genetic variants using the birthday model
[9]   A Covering Method for Detecting Genetic Associations between Rare Variants and Common Phenotypes [J].
Bhatia, Gaurav ;
Bansal, Vikas ;
Harismendy, Olivier ;
Schork, Nicholas J. ;
Topol, Eric J. ;
Frazer, Kelly ;
Bafna, Vineet .
PLOS COMPUTATIONAL BIOLOGY, 2010, 6 (10)
[10]   Identifying Mendelian disease genes with the Variant Effect Scoring Tool [J].
Carter, Hannah ;
Douville, Christopher ;
Stenson, Peter D. ;
Cooper, David N. ;
Karchin, Rachel .
BMC GENOMICS, 2013, 14