RNA-QC-chain: comprehensive and fast quality control for RNA-Seq data

被引:37
|
作者
Zhou, Qian [1 ,2 ]
Su, Xiaoquan [3 ,4 ,5 ]
Jing, Gongchao [3 ,4 ]
Chen, Songlin [1 ,2 ]
Ning, Kang [6 ]
机构
[1] Chinese Acad Fishery Sci, Key Lab Sustainable Dev Marine Fisheries, Minist Agr, Yellow Sea Fisheries Res Inst, Qingdao 266071, Shandong, Peoples R China
[2] Qingdao Natl Lab Marine Sci & Technol, Lab Marine Fisheries Sci & Food Prod Proc, Qingdao 266071, Shandong, Peoples R China
[3] Chinese Acad Sci, Qingdao Inst Bioenergy & Bioproc Technol, CAS Key Lab Biofuels, Shandong Key Lab Energy Genet, Qingdao 266101, Shandong, Peoples R China
[4] Chinese Acad Sci, Qingdao Inst Bioenergy & Bioproc Technol, Single Cell Ctr, Qingdao 266101, Shandong, Peoples R China
[5] Univ Chinese Acad Sci, Beijing 100049, Peoples R China
[6] Huazhong Univ Sci & Technol, Coll Life Sci & Technol, Hubei Key Lab Bioinformat & Mol Imaging, Minist Educ,Key Lab Mol Biophys,Dept Bioinformat, Wuhan 430074, Hubei, Peoples R China
来源
BMC GENOMICS | 2018年 / 19卷
基金
中国国家自然科学基金;
关键词
Quality control; RNA-Seq; Contamination identification; Alignment statistics; Parallel computing;
D O I
10.1186/s12864-018-4503-6
中图分类号
Q81 [生物工程学(生物技术)]; Q93 [微生物学];
学科分类号
071005 ; 0836 ; 090102 ; 100705 ;
摘要
Background: RNA-Seq has become one of the most widely used applications based on next-generation sequencing technology. However, raw RNA-Seq data may have quality issues, which can significantly distort analytical results and lead to erroneous conclusions. Therefore, the raw data must be subjected to vigorous quality control (QC) procedures before downstream analysis. Currently, an accurate and complete QC of RNA-Seq data requires of a suite of different QC tools used consecutively, which is inefficient in terms of usability, running time, file usage, and interpretability of the results. Results: We developed a comprehensive, fast and easy-to-use QC pipeline for RNA-Seq data, RNA-QC-Chain, which involves three steps: (1) sequencing-quality assessment and trimming; (2) internal (ribosomal RNAs) and external (reads from foreign species) contamination filtering; (3) alignment statistics reporting (such as read number, alignment coverage, sequencing depth and pair-end read mapping information). This package was developed based on our previously reported tool for general QC of next-generation sequencing (NGS) data called QC-Chain, with extensions specifically designed for RNA-Seq data. It has several features that are not available yet in other QC tools for RNA-Seq data, such as RNA sequence trimming, automatic rRNA detection and automatic contaminating species identification. The three QC steps can run either sequentially or independently, enabling RNA-QC-Chain as a comprehensive package with high flexibility and usability. Moreover, parallel computing and optimizations are embedded in most of the QC procedures, providing a superior efficiency. The performance of RNA-QC-Chain has been evaluated with different types of datasets, including an in-house sequencing data, a semi-simulated data, and two real datasets downloaded from public database. Comparisons of RNA-QC-Chain with other QC tools have manifested its superiorities in both function versatility and processing speed. Conclusions: We present here a tool, RNA-QC-Chain, which can be used to comprehensively resolve the quality control processes of RNA-Seq data effectively and efficiently.
引用
收藏
页数:10
相关论文
共 50 条
  • [21] Comprehensive evaluation of RNA-seq quantification methods for linearity
    Jin, Haijing
    Wan, Ying-Wooi
    Liu, Zhandong
    BMC BIOINFORMATICS, 2017, 18
  • [22] Comprehensive evaluation of RNA-seq quantification methods for linearity
    Haijing Jin
    Ying-Wooi Wan
    Zhandong Liu
    BMC Bioinformatics, 18
  • [23] A method for the extraction of high quality fungal RNA suitable for RNA-seq
    Cortes-Maldonado, Leyda
    Marcial-Quino, Jaime
    Gomez-Manzo, Saul
    Fierro, Francisco
    Tomasini, Araceli
    JOURNAL OF MICROBIOLOGICAL METHODS, 2020, 170
  • [24] RNA-seq library preparation for comprehensive transcriptome analysis in cancer cells
    Jaksik, Roman
    Drobna-Sledzinska, Monika
    Dawidowska, Malgorzata
    GENOMICS, 2021, 113 (06) : 4149 - 4162
  • [25] An Integrated Approach for RNA-seq Data Normalization
    Yang, Shengping
    Mercante, Donald E.
    Zhang, Kun
    Fang, Zhide
    CANCER INFORMATICS, 2016, 15 : 129 - 141
  • [26] eQTL Mapping Using RNA-seq Data
    Sun W.
    Hu Y.
    Statistics in Biosciences, 2013, 5 (1) : 198 - 219
  • [27] FastqPuri: high-performance preprocessing of RNA-seq data
    Paula Pérez-Rubio
    Claudio Lottaz
    Julia C. Engelmann
    BMC Bioinformatics, 20
  • [28] FastqPuri: high-performance preprocessing of RNA-seq data
    Perez-Rubio, Paula
    Lottaz, Claudio
    Engelmann, Julia C.
    BMC BIOINFORMATICS, 2019, 20 (1)
  • [29] Computational Considerations in Transcriptome Assemblies and Their Evaluation, using High Quality Human RNA-Seq data
    Ghaffari, Noushin
    Abante, Jordi
    Singh, Raminder
    Blood, Philip D.
    Johnson, Charles D.
    PROCEEDINGS OF XSEDE16: DIVERSITY, BIG DATA, AND SCIENCE AT SCALE, 2016,
  • [30] Comprehensive RNA-Seq Data Analysis Identifies Key mRNAs and lncRNAs in Atrial Fibrillation
    Wu, Dong-Mei
    Zhou, Zheng-Kun
    Fan, Shao-Hua
    Zheng, Zi-Hui
    Wen, Xin
    Han, Xin-Rui
    Wang, Shan
    Wang, Yong-Jian
    Zhang, Zi-Feng
    Shan, Qun
    Li, Meng-Qiu
    Hu, Bin
    Lu, Jun
    Chen, Gui-Quan
    Hong, Xiao-Wu
    Zheng, Yuan-Lin
    FRONTIERS IN GENETICS, 2019, 10