Genomics pipelines and data integration: challenges and opportunities in the research setting

被引:47
作者
Davis-Turak, Jeremy [1 ]
Courtney, Sean M. [2 ,3 ]
Hazard, E. Starr [2 ,4 ]
Glen, W. Bailey [2 ,3 ]
da Silveira, Willian A. [2 ,3 ]
Wesselman, Timothy [1 ]
Harbin, Larry P. [5 ]
Wolf, Bethany J. [5 ]
Chung, Dongjun [5 ]
Hardiman, Gary [2 ,5 ,6 ]
机构
[1] OnRamp Bioinformat Inc, San Diego, CA USA
[2] Med Univ South Carolina, MUSC Bioinformat, Ctr Genom Med, Charleston, SC 29425 USA
[3] Med Univ South Carolina, Dept Pathol & Lab Med, Charleston, SC USA
[4] Med Univ South Carolina, Lib Sci & Informat, Charleston, SC USA
[5] Med Univ South Carolina, Dept Publ Hlth Sci, Charleston, SC 29425 USA
[6] Med Univ South Carolina, Dept Med, Charleston, SC 29425 USA
基金
美国国家卫生研究院;
关键词
High throughput sequencing; bioinformatics pipelines; bioinformatics best practices; RNAseq; ExomeSeq; variant calling; reproducible computational research; genomic data management; analysis provenance; DIFFERENTIAL EXPRESSION ANALYSIS; RNA-SEQ; COMPREHENSIVE ANALYSIS; READ ALIGNMENT; CANCER; DISCOVERY; FRAMEWORK; GENE; TOOL; LANDSCAPE;
D O I
10.1080/14737159.2017.1282822
中图分类号
R36 [病理学];
学科分类号
100104 ;
摘要
Introduction: The emergence and mass utilization of high-throughput (HT) technologies, including sequencing technologies (genomics) and mass spectrometry (proteomics, metabolomics, lipids), has allowed geneticists, biologists, and biostatisticians to bridge the gap between genotype and phenotype on a massive scale. These new technologies have brought rapid advances in our understanding of cell biology, evolutionary history, microbial environments, and are increasingly providing new insights and applications towards clinical care and personalized medicine.Areas covered: The very success of this industry also translates into daunting big data challenges for researchers and institutions that extend beyond the traditional academic focus of algorithms and tools. The main obstacles revolve around analysis provenance, data management of massive datasets, ease of use of software, interpretability and reproducibility of results.Expert commentary: The authors review the challenges associated with implementing bioinformatics best practices in a large-scale setting, and highlight the opportunity for establishing bioinformatics pipelines that incorporate data tracking and auditing, enabling greater consistency and reproducibility for basic research, translational or clinical settings.
引用
收藏
页码:225 / 237
页数:13
相关论文
共 91 条
  • [21] A Comparison of Variant Calling Pipelines Using Genome in a Bottle as a Reference
    Cornish, Adam
    Guda, Chittibabu
    [J]. BIOMED RESEARCH INTERNATIONAL, 2015, 2015
  • [22] An Extensive Evaluation of Read Trimming Effects on Illumina NGS Data Analysis
    Del Fabbro, Cristian
    Scalabrin, Simone
    Morgante, Michele
    Giorgi, Federico M.
    [J]. PLOS ONE, 2013, 8 (12):
  • [23] A framework for variation discovery and genotyping using next-generation DNA sequencing data
    DePristo, Mark A.
    Banks, Eric
    Poplin, Ryan
    Garimella, Kiran V.
    Maguire, Jared R.
    Hartl, Christopher
    Philippakis, Anthony A.
    del Angel, Guillermo
    Rivas, Manuel A.
    Hanna, Matt
    McKenna, Aaron
    Fennell, Tim J.
    Kernytsky, Andrew M.
    Sivachenko, Andrey Y.
    Cibulskis, Kristian
    Gabriel, Stacey B.
    Altshuler, David
    Daly, Mark J.
    [J]. NATURE GENETICS, 2011, 43 (05) : 491 - +
  • [24] Cancer statistics for African Americans, 2016: Progress and opportunities in reducing racial disparities
    DeSantis, Carol E.
    Siegel, Rebecca L.
    Sauer, Ann Goding
    Miller, Kimberly D.
    Fedewa, Stacey A.
    Alcaraz, Kassandra I.
    Jemal, Ahmedin
    [J]. CA-A CANCER JOURNAL FOR CLINICIANS, 2016, 66 (04) : 290 - 308
  • [25] STAR: ultrafast universal RNA-seq aligner
    Dobin, Alexander
    Davis, Carrie A.
    Schlesinger, Felix
    Drenkow, Jorg
    Zaleski, Chris
    Jha, Sonali
    Batut, Philippe
    Chaisson, Mark
    Gingeras, Thomas R.
    [J]. BIOINFORMATICS, 2013, 29 (01) : 15 - 21
  • [26] Mechanisms Establishing TLR4-Responsive Activation States of Inflammatory Response Genes
    Escoubet-Lozach, Laure
    Benner, Christopher
    Kaikkonen, Minna U.
    Lozach, Jean
    Heinz, Sven
    Spann, Nathan J.
    Crotti, Andrea
    Stender, Josh
    Ghisletti, Serena
    Reichart, Donna
    Cheng, Christine S.
    Luna, Rosa
    Ludka, Colleen
    Sasik, Roman
    Garcia-Bassets, Ivan
    Hoffmann, Alexander
    Subramaniam, Shankar
    Hardiman, Gary
    Rosenfeld, Michael G.
    Glass, Christopher K.
    [J]. PLOS GENETICS, 2011, 7 (12):
  • [27] COSMIC: exploring the world's knowledge of somatic mutations in human cancer
    Forbes, Simon A.
    Beare, David
    Gunasekaran, Prasad
    Leung, Kenric
    Bindal, Nidhi
    Boutselakis, Harry
    Ding, Minjie
    Bamford, Sally
    Cole, Charlotte
    Ward, Sari
    Kok, Chai Yin
    Jia, Mingming
    De, Tisham
    Teague, Jon W.
    Stratton, Michael R.
    McDermott, Ultan
    Campbell, Peter J.
    [J]. NUCLEIC ACIDS RESEARCH, 2015, 43 (D1) : D805 - D811
  • [28] Garrison E., 2012, PREPRINT, DOI DOI 10.48550/ARXIV.1207.3907
  • [29] Gibbon G. A., 1996, Laboratory Automation and Information Management, V32, P1, DOI 10.1016/1381-141X(95)00024-K
  • [30] Toward a Shared Vision for Cancer Genomic Data
    Grossman, Robert L.
    Heath, Allison P.
    Ferretti, Vincent
    Varmus, Harold E.
    Lowy, Douglas R.
    Kibbe, Warren A.
    Staudt, Louis M.
    [J]. NEW ENGLAND JOURNAL OF MEDICINE, 2016, 375 (12) : 1109 - 1112