G4Boost: a machine learning-based tool for quadruplex identification and stability prediction

被引:14
|
作者
Cagirici, H. Busra [1 ]
Budak, Hikmet [2 ]
Sen, Taner Z. [1 ]
机构
[1] USDA ARS, Crop Improvement Genet Res Unit, Western Reg Res Ctr, 800 Buchanan St, Albany, CA 94710 USA
[2] Montana BioAgr Inc, Missoula, MT USA
关键词
G-quadruplex; Machine learning; Topology; Stability; Energy; Plants; Humans; RNA G-QUADRUPLEXES; SECONDARY STRUCTURE; WEB SERVER; DNA; PROMOTER; TRANSLATION; INHIBITION; PREVALENCE; TELOMERASE; SEQUENCE;
D O I
10.1186/s12859-022-04782-z
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Background G-quadruplexes (G4s), formed within guanine-rich nucleic acids, are secondary structures involved in important biological processes. Although every G4 motif has the potential to form a stable G4 structure, not every G4 motif would, and accurate energy-based methods are needed to assess their structural stability. Here, we present a decision tree-based prediction tool, G4Boost, to identify G4 motifs and predict their secondary structure folding probability and thermodynamic stability based on their sequences, nucleotide compositions, and estimated structural topologies. Results G4Boost predicted the quadruplex folding state with an accuracy greater then 93% and an F1-score of 0.96, and the folding energy with an RMSE of 4.28 and R-2 of 0.95 only by the means of sequence intrinsic feature. G4Boost was successfully applied and validated to predict the stability of experimentally-determined G4 structures, including for plants and humans. Conclusion G4Boost outperformed the three machine-learning based prediction tools, DeepG4, Quadron, and G4RNA Screener, in terms of both accuracy and F1-score, and can be highly useful for G4 prediction to understand gene regulation across species including plants and humans.
引用
收藏
页数:18
相关论文
共 50 条
  • [1] G4Boost: a machine learning-based tool for quadruplex identification and stability prediction
    H. Busra Cagirici
    Hikmet Budak
    Taner Z. Sen
    BMC Bioinformatics, 23
  • [2] Machine learning-based prediction of DNA G-quadruplex folding topology with G4ShapePredictor
    Liew, Donn
    Lim, Zi Way
    Yong, Ee Hou
    SCIENTIFIC REPORTS, 2024, 14 (01):
  • [3] A machine learning-based universal outbreak risk prediction tool
    Zhang, Tianyu
    Rabhi, Fethi
    Chen, Xin
    Paik, Hye-young
    Macintyre, Chandini Raina
    COMPUTERS IN BIOLOGY AND MEDICINE, 2024, 169
  • [4] Clinical evaluation of a machine learning-based dysphagia risk prediction tool
    Gugatschka, Markus
    Egger, Nina Maria
    Haspl, K.
    Hortobagyi, David
    Jauk, Stefanie
    Feiner, Marlies
    Kramer, Diether
    EUROPEAN ARCHIVES OF OTO-RHINO-LARYNGOLOGY, 2024, 281 (08) : 4379 - 4384
  • [5] Machine learning-based genetic feature identification and fatigue life prediction
    Zhou, Kun
    Sun, Xingyue
    Shi, Shouwen
    Song, Kai
    Chen, Xu
    FATIGUE & FRACTURE OF ENGINEERING MATERIALS & STRUCTURES, 2021, 44 (09) : 2524 - 2537
  • [6] A Machine Learning-Based Web Tool for the Severity Prediction of COVID-19
    Christodoulou, Avgi
    Katsarou, Martha-Spyridoula
    Emmanouil, Christina
    Gavrielatos, Marios
    Georgiou, Dimitrios
    Tsolakou, Annia
    Papasavva, Maria
    Economou, Vasiliki
    Nanou, Vasiliki
    Nikolopoulos, Ioannis
    Daganou, Maria
    Argyraki, Aikaterini
    Stefanidis, Evaggelos
    Metaxas, Gerasimos
    Panagiotou, Emmanouil
    Michalopoulos, Ioannis
    Drakoulis, Nikolaos
    BIOTECH, 2024, 13 (03):
  • [7] A Machine Learning-Based Severity Prediction Tool for the Michigan Neuropathy Screening Instrument
    Haque, Fahmida
    Reaz, Mamun B. I.
    Chowdhury, Muhammad E. H.
    bin Shapiai, Mohd Ibrahim
    Malik, Rayaz S. A.
    Alhatou, Mohammed
    Kobashi, Syoji
    Ara, Iffat
    Ali, Sawal H. M.
    Bakar, Ahmad A. A.
    Bhuiyan, Mohammad Arif Sobhan
    DIAGNOSTICS, 2023, 13 (02)
  • [8] Machine learning-based methods for MCS prediction in 5G networks
    Tsipi, Lefteris
    Karavolos, Michail
    Papaioannou, Grigorios
    Volakaki, Maria
    Vouyioukas, Demosthenes
    TELECOMMUNICATION SYSTEMS, 2024, 86 (04) : 705 - 728
  • [9] Machine Learning-Based Prediction of Stability in High-Entropy Nitride Ceramics
    Lin, Tianyu
    Wang, Ruolan
    Liu, Dazhi
    CRYSTALS, 2024, 14 (05)
  • [10] Machine learning-based cache miss prediction
    Jelacic, Edin
    Seceleanu, Cristina
    Xiong, Ning
    Backeman, Peter
    Yaghoobi, Sharifeh
    Seceleanu, Tiberiu
    INTERNATIONAL JOURNAL ON SOFTWARE TOOLS FOR TECHNOLOGY TRANSFER, 2025, : 53 - 80