Unified access to up-to-date residue-level annotations from UniProtKB and other biological databases for PDB data

被引:5
作者
Choudhary, Preeti [1 ]
Anyango, Stephen [1 ]
Berrisford, John [1 ,2 ]
Tolchard, James [1 ,3 ]
Varadi, Mihaly [1 ]
Velankar, Sameer [1 ]
机构
[1] European Bioinformat Inst EMBL EBI, European Mol Biol Lab, Prot Data Bank Europe, Wellcome Genome Campus, Cambridge CB10 1SD, Cambridgeshire, England
[2] AstraZeneca, Biomed Campus,1 Francis Crick Ave, Cambridge CB2 0AA, Cambridgeshire, England
[3] Claude Bernard Univ, F-69100 Villeurbanne, Lyon, France
关键词
TYROSINE-PHOSPHATASE; 1B; PROTEIN STRUCTURES; CRYSTAL-STRUCTURE; ENZYME; INHIBITORS; DISCOVERY; FAMILIES; STATES;
D O I
10.1038/s41597-023-02101-6
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
More than 61,000 proteins have up-to-date correspondence between their amino acid sequence (UniProtKB) and their 3D structures (PDB), enabled by the Structure Integration with Function, Taxonomy and Sequences (SIFTS) resource. SIFTS incorporates residue-level annotations from many other biological resources. SIFTS data is available in various formats like XML, CSV and TSV format or also accessible via the PDBe REST API but always maintained separately from the structure data (PDBx/mmCIF file) in the PDB archive. Here, we extended the wwPDB PDBx/mmCIF data dictionary with additional categories to accommodate SIFTS data and added the UniProtKB, Pfam, SCOP2, and CATH residue-level annotations directly into the PDBx/mmCIF files from the PDB archive. With the integrated UniProtKB annotations, these files now provide consistent numbering of residues in different PDB entries allowing easy comparison of structure models. The extended dictionary yields a more consistent, standardised metadata description without altering the core PDB information. This development enables up-to-date cross-reference information at the residue level resulting in better data interoperability, supporting improved data analysis and visualisation.
引用
收藏
页数:13
相关论文
共 77 条
[1]  
Agarwala R, 2018, NUCLEIC ACIDS RES, V46, pD8, DOI [10.1093/nar/gks1189, 10.1093/nar/gkx1095, 10.1093/nar/gkq1172]
[2]   Data growth and its impact on the SCOP database: new developments [J].
Andreeva, Antonina ;
Howorth, Dave ;
Chandonia, John-Marc ;
Brenner, Steven E. ;
Hubbard, Tim J. P. ;
Chothia, Cyrus ;
Murzin, Alexey G. .
NUCLEIC ACIDS RESEARCH, 2008, 36 :D419-D425
[3]   The SCOP database in 2020: expanded classification of representative family and superfamily domains of known protein structures [J].
Andreeva, Antonina ;
Kulesha, Eugene ;
Gough, Julian ;
Murzin, Alexey G. .
NUCLEIC ACIDS RESEARCH, 2020, 48 (D1) :D376-D382
[4]  
Andreeva Antonina, 2015, Curr Protoc Bioinformatics, V49, DOI 10.1002/0471250953.bi0126s49
[5]   FAIR principles for data stewardship [J].
不详 .
NATURE GENETICS, 2016, 48 (04) :343-343
[6]   PDBe: improved findability of macromolecular structure data in the PDB [J].
Armstrong, David R. ;
Berrisford, John M. ;
Conroy, Matthew J. ;
Gutmanas, Aleksandras ;
Anyango, Stephen ;
Choudhary, Preeti ;
Clark, Alice R. ;
Dana, Jose M. ;
Deshpande, Mandar ;
Dunlop, Roisin ;
Gane, Paul ;
Gaborova, Romana ;
Gupta, Deepti ;
Haslam, Pauline ;
Koca, Jaroslav ;
Mak, Lora ;
Mir, Saqib ;
Mukhopadhyay, Abhik ;
Nadzirin, Nurul ;
Nair, Sreenath ;
Paysan-Lafosse, Typhaine ;
Pravda, Lukas ;
Sehnal, David ;
Salih, Osman ;
Smart, Oliver ;
Tolchard, James ;
Varadi, Mihaly ;
Svobodova-Varekova, Radka ;
Zaki, Hossam ;
Kleywegt, Gerard J. ;
Velankar, Sameer .
NUCLEIC ACIDS RESEARCH, 2020, 48 (D1) :D335-D343
[7]   Domain insertions in protein structures [J].
Aroul-Selvam, R ;
Hubbard, T ;
Sasidharan, R .
JOURNAL OF MOLECULAR BIOLOGY, 2004, 338 (04) :633-641
[8]   Gene Ontology: tool for the unification of biology [J].
Ashburner, M ;
Ball, CA ;
Blake, JA ;
Botstein, D ;
Butler, H ;
Cherry, JM ;
Davis, AP ;
Dolinski, K ;
Dwight, SS ;
Eppig, JT ;
Harris, MA ;
Hill, DP ;
Issel-Tarver, L ;
Kasarskis, A ;
Lewis, S ;
Matese, JC ;
Richardson, JE ;
Ringwald, M ;
Rubin, GM ;
Sherlock, G .
NATURE GENETICS, 2000, 25 (01) :25-29
[9]   Engineered single-chain dimeric streptavidins with an unexpected strong preference for biotin-4-fluorescein [J].
Aslan, FM ;
Yu, Y ;
Mohr, SC ;
Cantor, CR .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 2005, 102 (24) :8507-8512
[10]   Accurate prediction of protein structures and interactions using a three-track neural network [J].
Baek, Minkyung ;
DiMaio, Frank ;
Anishchenko, Ivan ;
Dauparas, Justas ;
Ovchinnikov, Sergey ;
Lee, Gyu Rie ;
Wang, Jue ;
Cong, Qian ;
Kinch, Lisa N. ;
Schaeffer, R. Dustin ;
Millan, Claudia ;
Park, Hahnbeom ;
Adams, Carson ;
Glassman, Caleb R. ;
DeGiovanni, Andy ;
Pereira, Jose H. ;
Rodrigues, Andria V. ;
van Dijk, Alberdina A. ;
Ebrecht, Ana C. ;
Opperman, Diederik J. ;
Sagmeister, Theo ;
Buhlheller, Christoph ;
Pavkov-Keller, Tea ;
Rathinaswamy, Manoj K. ;
Dalwadi, Udit ;
Yip, Calvin K. ;
Burke, John E. ;
Garcia, K. Christopher ;
Grishin, Nick V. ;
Adams, Paul D. ;
Read, Randy J. ;
Baker, David .
SCIENCE, 2021, 373 (6557) :871-+