Predicting sub-cellular localization of tRNA synthetases from their primary structures

被引:5
作者
Panwar, Bharat [1 ]
Raghava, G. P. S. [1 ]
机构
[1] Inst Microbial Technol CSIR, Bioinformat Ctr, Sect 39A, Chandigarh, India
关键词
Mitochondrial tRNA synthetase; Support vector machine; Prediction; MARSpred; SUPPORT VECTOR MACHINE; AMINO-ACID; PROTEIN; MITOCHONDRIAL; CLASSIFICATION; GENE; IMPORT; TOOL; SVM;
D O I
10.1007/s00726-011-0872-8
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
Since endo-symbiotic events occur, all genes of mitochondrial aminoacyl tRNA synthetase (AARS) were lost or transferred from ancestral mitochondrial genome into the nucleus. The canonical pattern is that both cytosolic and mitochondrial AARSs coexist in the nuclear genome. In the present scenario all mitochondrial AARSs are nucleus-encoded, synthesized on cytosolic ribosomes and post-translationally imported from the cytosol into the mitochondria in eukaryotic cell. The site-based discrimination between similar types of enzymes is very challenging because they have almost same physico-chemical properties. It is very important to predict the sub-cellular location of AARSs, to understand the mitochondrial protein synthesis. We have analyzed and optimized the distinguishable patterns between cytosolic and mitochondrial AARSs. Firstly, support vector machines (SVM)-based modules have been developed using amino acid and dipeptide compositions and achieved Mathews correlation coefficient (MCC) of 0.82 and 0.73, respectively. Secondly, we have developed SVM modules using position-specific scoring matrix and achieved the maximum MCC of 0.78. Thirdly, we developed SVM modules using N-terminal, intermediate residues, C-terminal and split amino acid composition (SAAC) and achieved MCC of 0.82, 0.70, 0.39 and 0.86, respectively. Finally, a SVM module was developed using selected attributes of split amino acid composition (SA-SAAC) approach and achieved MCC of 0.92 with an accuracy of 96.00%. All modules were trained and tested on a non-redundant data set and evaluated using fivefold cross-validation technique. On the independent data sets, SA-SAAC based prediction model achieved MCC of 0.95 with an accuracy of 97.77%. The web-server 'MARSpred' based on above study is available at http://www.imtech.res.in/raghava/marspred/.
引用
收藏
页码:1703 / 1713
页数:11
相关论文
共 35 条
[31]   Development of tRNA synthetases and connection to genetic code and disease [J].
Schimmel, Paul .
PROTEIN SCIENCE, 2008, 17 (10) :1643-1652
[32]   Predotar:: A tool for rapidly screening proteomes for N-terminal targeting sequences [J].
Small, I ;
Peeters, N ;
Legeai, F ;
Lurin, C .
PROTEOMICS, 2004, 4 (06) :1581-1590
[33]   Evidence that the mitochondrial leucyl tRNA synthetase (LARS2) gene represents a novel type 2 diabetes susceptibility gene [J].
t Hart, LM ;
Hansen, T ;
Rietveld, I ;
Dekker, JM ;
Nijpels, G ;
Janssen, GMC ;
Arp, PA ;
Uitterlinden, AG ;
Jorgensen, T ;
Borch-Johnsen, K ;
Pols, HAP ;
Pedersen, O ;
van Duijn, CM ;
Heine, RJ ;
Maassen, JA .
DIABETES, 2005, 54 (06) :1892-1895
[34]   The mitochondrial genome of Arabidopsis thaliana contains 57 genes in 366,924 nucleotides [J].
Unseld, M ;
Marienfeld, JR ;
Brandt, P ;
Brennicke, A .
NATURE GENETICS, 1997, 15 (01) :57-61
[35]   An overview of statistical learning theory [J].
Vapnik, VN .
IEEE TRANSACTIONS ON NEURAL NETWORKS, 1999, 10 (05) :988-999