Denoising the Denoisers: an independent evaluation of microbiome sequence error-correction approaches

被引:219
作者
Nearing, Jacob T. [1 ]
Douglas, Gavin M. [1 ]
Comeau, Andre M. [2 ]
Langille, Morgan G., I [1 ,3 ]
机构
[1] Dalhousie Univ, Dept Microbiol & Immunol, Halifax, NS, Canada
[2] Dalhousie Univ, Integrated Microbiome Resource, Halifax, NS, Canada
[3] Dalhousie Univ, Dept Pharmacol, Halifax, NS, Canada
来源
PEERJ | 2018年 / 6卷
基金
加拿大自然科学与工程研究理事会;
关键词
Microbiome; Denoising tools; Comparison; DADA2; Deblur; UNOISE3; Mock community; BACTERIAL; DATABASE; SEARCH;
D O I
10.7717/peerj.5364
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
High-depth sequencing of universal marker genes such as the 16S rRNA gene is a common strategy to profile microbial communities. Traditionally, sequence reads are clustered into operational taxonomic units (OTUs) at a defined identity threshold to avoid sequencing errors generating spurious taxonomic units. However, there have been numerous bioinformatic packages recently released that attempt to correct sequencing errors to determine real biological sequences at single nucleotide resolution by generating amplicon sequence variants (ASVs). As more researchers begin to use high resolution ASVs, there is a need for an in-depth and unbiased comparison of these novel "denoising" pipelines. In this study, we conduct a thorough comparison of three of the most widely-used denoising packages (DADA2, UNOISE3, and Deblur) as well as an open-reference 97% OTU clustering pipeline on mock, soil, and host-associated communities. We found from the mock community analyses that although they produced similar microbial compositions based on relative abundance, the approaches identified vastly different numbers of ASVs that significantly impact alpha diversity metrics. Our analysis on real datasets using recommended settings for each denoising pipeline also showed that the three packages were consistent in their per-sample compositions, resulting in only minor differences based on weighted UniFrac and Bray-Curtis dissimilarity. DADA2 tended to find more ASVs than the other two denoising pipelines when analyzing both the real soil data and two other host-associated datasets, suggesting that it could be better at finding rare organisms, but at the expense of possible false positives. The open-reference OTU clustering approach identified considerably more OTUs in comparison to the number of ASVs from the denoising pipelines in all datasets tested. The three denoising approaches were significantly different in their run times, with UNOISE3 running greater than 1,200 and 15 times faster than DADA2 and Deblur, respectively. Our findings indicate that, although all pipelines result in similar general community structure, the number of ASVs/OTUs and resulting alpha-diversity metrics varies considerably and should be considered when attempting to identify rare organisms from possible background noise.
引用
收藏
页数:22
相关论文
共 31 条
[11]   QIIME allows analysis of high-throughput community sequencing data [J].
Caporaso, J. Gregory ;
Kuczynski, Justin ;
Stombaugh, Jesse ;
Bittinger, Kyle ;
Bushman, Frederic D. ;
Costello, Elizabeth K. ;
Fierer, Noah ;
Pena, Antonio Gonzalez ;
Goodrich, Julia K. ;
Gordon, Jeffrey I. ;
Huttley, Gavin A. ;
Kelley, Scott T. ;
Knights, Dan ;
Koenig, Jeremy E. ;
Ley, Ruth E. ;
Lozupone, Catherine A. ;
McDonald, Daniel ;
Muegge, Brian D. ;
Pirrung, Meg ;
Reeder, Jens ;
Sevinsky, Joel R. ;
Tumbaugh, Peter J. ;
Walters, William A. ;
Widmann, Jeremy ;
Yatsunenko, Tanya ;
Zaneveld, Jesse ;
Knight, Rob .
NATURE METHODS, 2010, 7 (05) :335-336
[12]   Ribosomal Database Project: data and tools for high throughput rRNA analysis [J].
Cole, James R. ;
Wang, Qiong ;
Fish, Jordan A. ;
Chai, Benli ;
McGarrell, Donna M. ;
Sun, Yanni ;
Brown, C. Titus ;
Porras-Alfaro, Andrea ;
Kuske, Cheryl R. ;
Tiedje, James M. .
NUCLEIC ACIDS RESEARCH, 2014, 42 (D1) :D633-D642
[13]   Microbiome Helper: a Custom and Streamlined Workflow for Microbiome Research [J].
Comeau, Andre M. ;
Douglas, Gavin M. ;
Langille, Morgan G. I. .
MSYSTEMS, 2017, 2 (01)
[14]   Greengenes, a chimera-checked 16S rRNA gene database and workbench compatible with ARB [J].
DeSantis, T. Z. ;
Hugenholtz, P. ;
Larsen, N. ;
Rojas, M. ;
Brodie, E. L. ;
Keller, K. ;
Huber, T. ;
Dalevi, D. ;
Hu, P. ;
Andersen, G. L. .
APPLIED AND ENVIRONMENTAL MICROBIOLOGY, 2006, 72 (07) :5069-5072
[15]   Multi-omics differentially classify disease state and treatment outcome in pediatric Crohn's disease [J].
Douglas, Gavin M. ;
Hansen, Richard ;
Jones, Casey M. A. ;
Dunn, Katherine A. ;
Comeau, Andre M. ;
Bielawski, Joseph P. ;
Tayler, Rachel ;
El-Omar, Emad M. ;
Russell, Richard K. ;
Hold, Georgina L. ;
Langille, Morgan G. I. ;
Van Limbergen, Johan .
MICROBIOME, 2018, 6
[16]   Accuracy of microbial community diversity estimated by closed- and open-reference OTUs [J].
Edgar, Robert C. .
PEERJ, 2017, 5
[17]   UCHIME improves sensitivity and speed of chimera detection [J].
Edgar, Robert C. ;
Haas, Brian J. ;
Clemente, Jose C. ;
Quince, Christopher ;
Knight, Rob .
BIOINFORMATICS, 2011, 27 (16) :2194-2200
[18]   Search and clustering orders of magnitude faster than BLAST [J].
Edgar, Robert C. .
BIOINFORMATICS, 2010, 26 (19) :2460-2461
[19]   The diversity and biogeography of soil bacterial communities [J].
Fierer, N ;
Jackson, RB .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 2006, 103 (03) :626-631
[20]   Towards a unified paradigm for sequence-based identification of fungi [J].
Koljalg, Urmas ;
Nilsson, R. Henrik ;
Abarenkov, Kessy ;
Tedersoo, Leho ;
Taylor, Andy F. S. ;
Bahram, Mohammad ;
Bates, Scott T. ;
Bruns, Thomas D. ;
Bengtsson-Palme, Johan ;
Callaghan, Tony M. ;
Douglas, Brian ;
Drenkhan, Tiia ;
Eberhardt, Ursula ;
Duenas, Margarita ;
Grebenc, Tine ;
Griffith, Gareth W. ;
Hartmann, Martin ;
Kirk, Paul M. ;
Kohout, Petr ;
Larsson, Ellen ;
Lindahl, Bjoern D. ;
Luecking, Robert ;
Martin, Maria P. ;
Matheny, P. Brandon ;
Nguyen, Nhu H. ;
Niskanen, Tuula ;
Oja, Jane ;
Peay, Kabir G. ;
Peintner, Ursula ;
Peterson, Marko ;
Poldmaa, Kadri ;
Saag, Lauri ;
Saar, Irja ;
Schüessler, Arthur ;
Scott, James A. ;
Senes, Carolina ;
Smith, Matthew E. ;
Suija, Ave ;
Taylor, D. Lee ;
Telleria, M. Teresa ;
Weiss, Michael ;
Larsson, Karl-Henrik .
MOLECULAR ECOLOGY, 2013, 22 (21) :5271-5277