Marker imputation efficiency for genotyping-by-sequencing data in rice (Oryza sativa) and alfalfa (Medicago sativa)

被引:39
作者
Nazzicari, Nelson [1 ]
Biscarini, Filippo [2 ]
Cozzi, Paolo [2 ]
Brummer, E. Charles [3 ]
Annicchiarico, Paolo [2 ]
机构
[1] Res Ctr Fodder Crops & Dairy Prod, Council Agr Res & Econ CREA, Lodi, Italy
[2] Fdn Parco Tecnol Padano, Dipartimento Bioinformat, Lodi, Italy
[3] Univ Calif Davis, Dept Plant Sci, Davis, CA 95616 USA
关键词
SNP; Genotyping by sequencing (GBS); K-nearest neighbors imputation (KNNI); Random Forest imputation (RFI); Singular value decomposition imputation (SVDI); Beagle; FILLIN; Alfalfa; Rice; Imputation; Reference genome; GENOMIC SELECTION; READ ALIGNMENT; LINKAGE MAP; ASSOCIATION; POPULATIONS; ACCURACY;
D O I
10.1007/s11032-016-0490-y
中图分类号
S3 [农学(农艺学)];
学科分类号
0901 ;
摘要
Genotyping-by-sequencing (GBS) is a rapid and cost-effective genome-wide genotyping technique applicable whether a reference genome is available or not. Due to the cost-coverage trade-off, however, GBS typically produces large amounts of missing marker genotypes, whose imputation becomes therefore both challenging and critical for later analyses. In this work, the performance of four general imputation methods (K-nearest neighbors, Random Forest, singular value decomposition, and mean value) and two genotype-specific methods ("Beagle" and FILLIN) was measured on GBS data from alfalfa (Medicago sativa L., autotetraploid, heterozygous, without reference genome) and rice (Oryza sativa L., diploid, 100 % homozygous, with reference genome). Alfalfa SNP were aligned on the genome of the closely related species Medicago truncatula L.. Benchmarks consisted in progressive data filtering for marker call rate (up to 70 %) and increasing proportions (up to 20 %) of known genotypes masked for imputation. The relative performance was measured as the total proportion of correctly imputed genotypes, globally and within each genotype class (two homozygotes in rice, two homozygotes and one heterozygote in alfalfa). We found that imputation accuracy was robust to increasing missing rates, and consistently higher in rice than in alfalfa. Accuracy was as high as 90-100 % for the major (most frequent) homozygous genotype, but dropped to 80-90 %(rice) and below 30 %(alfalfa) in the minor homozygous genotype. Beagle was the best performing method, both accuracy-and time-wise, in rice. In alfalfa, KNNI and RFI gave the highest accuracies, but KNNI was much faster.
引用
收藏
页数:16
相关论文
共 46 条
[1]   An integrated map of genetic variation from 1,092 human genomes [J].
Altshuler, David M. ;
Durbin, Richard M. ;
Abecasis, Goncalo R. ;
Bentley, David R. ;
Chakravarti, Aravinda ;
Clark, Andrew G. ;
Donnelly, Peter ;
Eichler, Evan E. ;
Flicek, Paul ;
Gabriel, Stacey B. ;
Gibbs, Richard A. ;
Green, Eric D. ;
Hurles, Matthew E. ;
Knoppers, Bartha M. ;
Korbel, Jan O. ;
Lander, Eric S. ;
Lee, Charles ;
Lehrach, Hans ;
Mardis, Elaine R. ;
Marth, Gabor T. ;
McVean, Gil A. ;
Nickerson, Deborah A. ;
Schmidt, Jeanette P. ;
Sherry, Stephen T. ;
Wang, Jun ;
Wilson, Richard K. ;
Gibbs, Richard A. ;
Dinh, Huyen ;
Kovar, Christie ;
Lee, Sandra ;
Lewis, Lora ;
Muzny, Donna ;
Reid, Jeff ;
Wang, Min ;
Wang, Jun ;
Fang, Xiaodong ;
Guo, Xiaosen ;
Jian, Min ;
Jiang, Hui ;
Jin, Xin ;
Li, Guoqing ;
Li, Jingxiang ;
Li, Yingrui ;
Li, Zhuo ;
Liu, Xiao ;
Lu, Yao ;
Ma, Xuedi ;
Su, Zhe ;
Tai, Shuaishuai ;
Tang, Meifang .
NATURE, 2012, 491 (7422) :56-65
[2]   Accuracy of genomic selection for alfalfa biomass yield in different reference populations [J].
Annicchiarico, Paolo ;
Nazzicari, Nelson ;
Li, Xuehui ;
Wei, Yanling ;
Pecetti, Luciano ;
Brummer, E. Charles .
BMC GENOMICS, 2015, 16
[3]   GenABEL: an R library for genome-wide association analysis [J].
Aulchenko, Yurii S. ;
Ripke, Stephan ;
Isaacs, Aaron ;
Van Duijn, Cornelia M. .
BIOINFORMATICS, 2007, 23 (10) :1294-1296
[4]   DYNAMIC PROGRAMMING [J].
BELLMAN, R .
SCIENCE, 1966, 153 (3731) :34-&
[5]   Genome-enabled predictions for binomial traits in sugar beet populations [J].
Biscarini, Filippo ;
Stevanato, Piergiorgio ;
Broccanello, Chiara ;
Stella, Alessandra ;
Saccomani, Massimo .
BMC GENETICS, 2014, 15
[6]   Random forests [J].
Breiman, L .
MACHINE LEARNING, 2001, 45 (01) :5-32
[7]   Short communication: Genotype imputation within and across Nordic cattle breeds [J].
Brondum, R. F. ;
Ma, P. ;
Lund, M. S. ;
Su, G. .
JOURNAL OF DAIRY SCIENCE, 2012, 95 (11) :6795-6800
[8]   Rapid and accurate haplotype phasing and missing-data inference for whole-genome association studies by use of localized haplotype clustering [J].
Browning, Sharon R. ;
Browning, Brian L. .
AMERICAN JOURNAL OF HUMAN GENETICS, 2007, 81 (05) :1084-1097
[9]  
Browningr B, 2011, BEAGLE 3 3 2
[10]   Genomic Prediction in Maize Breeding Populations with Genotyping-by-Sequencing [J].
Crossa, Jose ;
Beyene, Yoseph ;
Kassa, Semagn ;
Perez, Paulino ;
Hickey, John M. ;
Chen, Charles ;
de los Campos, Gustavo ;
Burgueno, Juan ;
Windhausen, Vanessa S. ;
Buckler, Ed ;
Jannink, Jean-Luc ;
Lopez Cruz, Marco A. ;
Babu, Raman .
G3-GENES GENOMES GENETICS, 2013, 3 (11) :1903-1926