RNA-QC-chain: comprehensive and fast quality control for RNA-Seq data

被引:37
|
作者
Zhou, Qian [1 ,2 ]
Su, Xiaoquan [3 ,4 ,5 ]
Jing, Gongchao [3 ,4 ]
Chen, Songlin [1 ,2 ]
Ning, Kang [6 ]
机构
[1] Chinese Acad Fishery Sci, Key Lab Sustainable Dev Marine Fisheries, Minist Agr, Yellow Sea Fisheries Res Inst, Qingdao 266071, Shandong, Peoples R China
[2] Qingdao Natl Lab Marine Sci & Technol, Lab Marine Fisheries Sci & Food Prod Proc, Qingdao 266071, Shandong, Peoples R China
[3] Chinese Acad Sci, Qingdao Inst Bioenergy & Bioproc Technol, CAS Key Lab Biofuels, Shandong Key Lab Energy Genet, Qingdao 266101, Shandong, Peoples R China
[4] Chinese Acad Sci, Qingdao Inst Bioenergy & Bioproc Technol, Single Cell Ctr, Qingdao 266101, Shandong, Peoples R China
[5] Univ Chinese Acad Sci, Beijing 100049, Peoples R China
[6] Huazhong Univ Sci & Technol, Coll Life Sci & Technol, Hubei Key Lab Bioinformat & Mol Imaging, Minist Educ,Key Lab Mol Biophys,Dept Bioinformat, Wuhan 430074, Hubei, Peoples R China
来源
BMC GENOMICS | 2018年 / 19卷
基金
中国国家自然科学基金;
关键词
Quality control; RNA-Seq; Contamination identification; Alignment statistics; Parallel computing;
D O I
10.1186/s12864-018-4503-6
中图分类号
Q81 [生物工程学(生物技术)]; Q93 [微生物学];
学科分类号
071005 ; 0836 ; 090102 ; 100705 ;
摘要
Background: RNA-Seq has become one of the most widely used applications based on next-generation sequencing technology. However, raw RNA-Seq data may have quality issues, which can significantly distort analytical results and lead to erroneous conclusions. Therefore, the raw data must be subjected to vigorous quality control (QC) procedures before downstream analysis. Currently, an accurate and complete QC of RNA-Seq data requires of a suite of different QC tools used consecutively, which is inefficient in terms of usability, running time, file usage, and interpretability of the results. Results: We developed a comprehensive, fast and easy-to-use QC pipeline for RNA-Seq data, RNA-QC-Chain, which involves three steps: (1) sequencing-quality assessment and trimming; (2) internal (ribosomal RNAs) and external (reads from foreign species) contamination filtering; (3) alignment statistics reporting (such as read number, alignment coverage, sequencing depth and pair-end read mapping information). This package was developed based on our previously reported tool for general QC of next-generation sequencing (NGS) data called QC-Chain, with extensions specifically designed for RNA-Seq data. It has several features that are not available yet in other QC tools for RNA-Seq data, such as RNA sequence trimming, automatic rRNA detection and automatic contaminating species identification. The three QC steps can run either sequentially or independently, enabling RNA-QC-Chain as a comprehensive package with high flexibility and usability. Moreover, parallel computing and optimizations are embedded in most of the QC procedures, providing a superior efficiency. The performance of RNA-QC-Chain has been evaluated with different types of datasets, including an in-house sequencing data, a semi-simulated data, and two real datasets downloaded from public database. Comparisons of RNA-QC-Chain with other QC tools have manifested its superiorities in both function versatility and processing speed. Conclusions: We present here a tool, RNA-QC-Chain, which can be used to comprehensively resolve the quality control processes of RNA-Seq data effectively and efficiently.
引用
收藏
页数:10
相关论文
共 50 条
  • [31] Modeling RNA degradation for RNA-Seq with applications
    Wan, Lin
    Yan, Xiting
    Chen, Ting
    Sun, Fengzhu
    BIOSTATISTICS, 2012, 13 (04) : 734 - 747
  • [32] Physiological RNA dynamics in RNA-Seq analysis
    Xu, Zhongneng
    Asakawa, Shuichi
    BRIEFINGS IN BIOINFORMATICS, 2019, 20 (05) : 1725 - 1733
  • [33] Quality Control Metrics for Extraction-Free Targeted RNA-Seq Under a Compositional Framework
    LaRoche, Dominic
    Billheimer, Dean
    Michels, Kurt
    LaFleur, Bonnie
    PHARMACEUTICAL STATISTICS (MBSW 39), 2019, 218 : 299 - 314
  • [34] SERE: Single-parameter quality control and sample comparison for RNA-Seq
    Stefan K Schulze
    Rahul Kanwar
    Meike Gölzenleuchter
    Terry M Therneau
    Andreas S Beutler
    BMC Genomics, 13
  • [35] SERE: Single-parameter quality control and sample comparison for RNA-Seq
    Schulze, Stefan K.
    Kanwar, Rahul
    Goelzenleuchter, Meike
    Therneau, Terry M.
    Beutler, Andreas S.
    BMC GENOMICS, 2012, 13 : 1 - 9
  • [36] The application of RNA-seq to the comprehensive analysis of plant mitochondrial transcriptomes
    Stone, James D.
    Storchova, Helena
    MOLECULAR GENETICS AND GENOMICS, 2015, 290 (01) : 1 - 9
  • [37] Influence of RNA extraction methods and library selection schemes on RNA-seq data
    Sultan, Marc
    Amstislavskiy, Vyacheslav
    Risch, Thomas
    Schuette, Moritz
    Doekel, Simon
    Ralser, Meryem
    Balzereit, Daniela
    Lehrach, Hans
    Yaspo, Marie-Laure
    BMC GENOMICS, 2014, 15
  • [38] Detection of generic differential RNA processing events from RNA-seq data
    Tran, Van Du T.
    Souiai, Oussema
    Romero-Barrios, Natali
    Crespi, Martin
    Gautheret, Daniel
    RNA BIOLOGY, 2016, 13 (01) : 59 - 67
  • [39] Plant Public RNA-seq Database: a comprehensive online database for expression analysis of ∼45 000 plant public RNA-Seq libraries
    Yu, Yiming
    Zhang, Hong
    Long, Yanping
    Shu, Yi
    Zhai, Jixian
    PLANT BIOTECHNOLOGY JOURNAL, 2022, 20 (05) : 806 - 808
  • [40] A Framework for Comparison and Assessment of Synthetic RNA-Seq Data
    Shakola, Felitsiya
    Palejev, Dean
    Ivanov, Ivan
    GENES, 2022, 13 (12)