rSeqTU-A Machine-Learning Based R Package for Prediction of Bacteria Transcription Units

被引:5
|
作者
Niu, Sheng-Yong [1 ]
Liu, Binqiang [2 ]
Ma, Qin [3 ]
Chou, Wen-Chi [4 ]
机构
[1] Univ Calif San Diego, Dept Comp Sci & Engn, La Jolla, CA 92093 USA
[2] Shandong Univ, Sch Math, Jinan, Shandong, Peoples R China
[3] Ohio State Univ, Coll Med, Biomed Informat, Columbus, OH 43210 USA
[4] Broad Inst MIT & Harvard, Infect Dis & Microbiome Program, Cambridge, MA 02142 USA
基金
美国国家科学基金会;
关键词
machine learning; bacteria; transcription unit; R package; transcriptome;
D O I
10.3389/fgene.2019.00374
中图分类号
Q3 [遗传学];
学科分类号
071007 ; 090102 ;
摘要
A transcription unit (TU) is composed of one or multiple adjacent genes on the same strand that are co-transcribed in mostly prokaryotes. Accurate identification of TUs is a crucial first step to delineate the transcriptional regulatory networks and elucidate the dynamic regulatory mechanisms encoded in various prokaryotic genomes. Many genomic features, for example, gene intergenic distance, and transcriptomic features including continuous and stable RNA-seq reads count signals, have been collected from a large amount of experimental data and integrated into classification techniques to computationally predict genome-wide TUs. Although some tools and web servers are able to predict TUs based on bacterial RNA-seq data and genome sequences, there is a need to have an improved machine learning prediction approach and a better comprehensive pipeline handling QC, TU prediction, and TU visualization. To enable users to efficiently perform TU identification on their local computers or high-performance clusters and provide a more accurate prediction, we develop an R package, named rSeqTU. rSeqTU uses a random forest algorithm to select essential features describing TUs and then uses support vector machine (SVM) to build TU prediction models. rSeqTU (available at https://s18692001.githubio/rSeqTU/) has six computational functionalities including read quality control, read mapping, training set generation, random forest-based feature selection, TU prediction, and TU visualization.
引用
收藏
页数:6
相关论文
共 50 条
  • [41] Machine-Learning Based Determination of Gait Events from Foot-Mounted Inertial Units
    Zago, Matteo
    Tarabini, Marco
    Delfino Spiga, Martina
    Ferrario, Cristina
    Bertozzi, Filippo
    Sforza, Chiarella
    Galli, Manuela
    SENSORS, 2021, 21 (03) : 1 - 13
  • [42] Chemprop: A Machine Learning Package for Chemical Property Prediction
    Heid, Esther
    Greenman, Kevin P.
    Chung, Yunsie
    Li, Shih-Cheng
    Graff, David E.
    Vermeire, Florence H.
    Wu, Haoyang
    Green, William H.
    Mcgill, Charles J.
    JOURNAL OF CHEMICAL INFORMATION AND MODELING, 2023, 64 (01) : 9 - 17
  • [43] assignPOP: An R package for population assignment using genetic, non-genetic, or integrated data in a machine-learning framework
    Chen, Kuan-Yu
    Marschall, Elizabeth A.
    Sovic, Michael G.
    Fries, Anthony C.
    Gibbs, H. Lisle
    Ludsin, Stuart A.
    METHODS IN ECOLOGY AND EVOLUTION, 2018, 9 (02): : 439 - 446
  • [44] ExplaineR: an R package to explain machine learning models
    Zargari Marandi, Ramtin
    BIOINFORMATICS ADVANCES, 2024, 4 (01):
  • [45] Predicting Vehicles' Positions using Roadside Units: a Machine-Learning Approach
    Sangare, Mamoudou
    Banerjee, Soumya
    Muhlethaler, Paul
    Bouzefrane, Samia
    2018 IEEE CONFERENCE ON STANDARDS FOR COMMUNICATIONS AND NETWORKING (IEEE CSCN), 2018,
  • [46] Machine-learning for the prediction of one-year seizure recurrence based on routine electroencephalography
    Émile Lemoine
    Denahin Toffa
    Geneviève Pelletier-Mc Duff
    An Qi Xu
    Mezen Jemel
    Jean-Daniel Tessier
    Frédéric Lesage
    Dang K. Nguyen
    Elie Bou Assi
    Scientific Reports, 13
  • [47] Structure-based prediction of BRAF mutation classes using machine-learning approaches
    Fanny S. Krebs
    Christian Britschgi
    Sylvain Pradervand
    Rita Achermann
    Petros Tsantoulis
    Simon Haefliger
    Andreas Wicki
    Olivier Michielin
    Vincent Zoete
    Scientific Reports, 12
  • [48] Prediction of future cognitive impairment among the community elderly: a machine-learning based approach
    Na, K. S.
    EUROPEAN PSYCHIATRY, 2019, 56 : S431 - S431
  • [49] Machine-learning based prediction of injection rate and solenoid voltage characteristics in GDI injectors
    Oh, Heechang
    Hwang, Joonsik
    Pickett, Lyle M.
    Han, Donghee
    FUEL, 2022, 311
  • [50] Landslide Susceptibility Prediction Considering Regional Soil Erosion Based on Machine-Learning Models
    Huang, Faming
    Chen, Jiawu
    Du, Zhen
    Yao, Chi
    Huang, Jinsong
    Jiang, Qinghui
    Chang, Zhilu
    Li, Shu
    ISPRS INTERNATIONAL JOURNAL OF GEO-INFORMATION, 2020, 9 (06)