Remote homology and the functions of metagenomic dark matter

被引:27
作者
Lobb, Briallen [1 ]
Kurtz, Daniel A. [1 ]
Moreno-Hagelsieb, Gabriel [2 ]
Doxey, Andrew C. [1 ]
机构
[1] Univ Waterloo, Dept Biol, Waterloo, ON N2L 3G1, Canada
[2] Wilfrid Laurier Univ, Dept Biol, Waterloo, ON N2L 3C5, Canada
关键词
ESCHERICHIA-COLI; PROTEIN; ORFANS; ORIGIN; GENES; EVOLUTION; INSIGHTS; ORPHANS; FAMILY; VIEW;
D O I
10.3389/fgene.2015.00234
中图分类号
Q3 [遗传学];
学科分类号
071007 ; 090102 ;
摘要
Predicted open reading frames (ORFs) that lack detectable homology to known proteins are termed ORFans. Despite their prevalence in metagenomes, the extent to which ORFans encode real proteins, the degree to which they can be annotated, and their functional contributions, remain unclear. To gain insights into these questions, we applied sensitive remote-homology detection methods to functionally analyze ORFans from soil, marine, and human gut metagenome collections. ORFans were identified, clustered into sequence families, and annotated through profile-profile comparison to proteins of known structure. We found that a considerable number of metagenomic ORFans (73,896 of 484,121, 15.3%) exhibit significant remote homology to structurally characterized proteins, providing a means for ORFan functional profiling. The extent of detected remote homology far exceeds that obtained for artificial protein families (1.4%). As expected for real genes, the predicted functions of ORFans are significantly similar to the functions of their gene neighbors (p < 0.001). Compared to the functional profiles predicted through standard homology searches, ORFans show biologically intriguing differences. Many ORFan-enriched functions are virus-related and tend to reflect biological processes associated with extreme sequence diversity. Each environment also possesses a large number of unique ORFan families and functions, including some known to play important community roles such as gut microbial polysaccharide digestion. Lastly, ORFans are a valuable resource for finding novel enzymes of interest, as we demonstrate through the identification of hundreds of novel ORFan metalloproteases that all possess a signature catalytic motif despite a general lack of similarity to known proteins. Our ORFan functional predictions are a valuable resource for discovering novel protein families and exploring the boundaries of protein sequence space. All remote homology predictions are available at http://doxey.uwaterloo.ca/ORFans.
引用
收藏
页数:12
相关论文
共 68 条
[11]   Structural motif screening reveals a novel, conserved carbohydrate-binding surface in the pathogenesis-related protein PR-5d [J].
Doxey, Andrew C. ;
Cheng, Zhenyu ;
Moffatt, Barbara A. ;
McConkey, Brendan J. .
BMC STRUCTURAL BIOLOGY, 2010, 10
[12]   Insights into the evolutionary origins of clostridial neurotoxins from analysis of the Clostridium botulinum strain A neurotoxin gene cluster [J].
Doxey, Andrew C. ;
Lynch, Michael D. J. ;
Mueller, Kirsten M. ;
Meiering, Elizabeth M. ;
McConkey, Brendan J. .
BMC EVOLUTIONARY BIOLOGY, 2008, 8 (1)
[13]   Bacterial collagenases - A review [J].
Duarte, Ana Sofia ;
Correia, Antonio ;
Esteves, Ana Cristina .
CRITICAL REVIEWS IN MICROBIOLOGY, 2016, 42 (01) :106-126
[14]   Analysis of bacterial community structure in sulfurous-oil-containing soils and detection of species carrying dibenzothiophene desulfurization (dsz) genes [J].
Duarte, GF ;
Rosado, AS ;
Seldin, L ;
de Araujo, W ;
van Elsas, JD .
APPLIED AND ENVIRONMENTAL MICROBIOLOGY, 2001, 67 (03) :1052-1062
[15]   The yeast genome project: What did we learn? [J].
Dujon, B .
TRENDS IN GENETICS, 1996, 12 (07) :263-270
[16]  
Fastrez J, 1996, EXS, V75, P35
[17]   Pfam: the protein families database [J].
Finn, Robert D. ;
Bateman, Alex ;
Clements, Jody ;
Coggill, Penelope ;
Eberhardt, Ruth Y. ;
Eddy, Sean R. ;
Heger, Andreas ;
Hetherington, Kirstie ;
Holm, Liisa ;
Mistry, Jaina ;
Sonnhammer, Erik L. L. ;
Tate, John ;
Punta, Marco .
NUCLEIC ACIDS RESEARCH, 2014, 42 (D1) :D222-D230
[18]   Polysaccharide utilization by gut bacteria: potential for new insights from genomic analysis [J].
Flint, Harry J. ;
Bayer, Edward A. ;
Rincon, Marco T. ;
Lamed, Raphael ;
White, Bryan A. .
NATURE REVIEWS MICROBIOLOGY, 2008, 6 (02) :121-131
[19]   Who's your neighbor? New computational approaches for functional genomics [J].
Galperin, MY ;
Koonin, EV .
NATURE BIOTECHNOLOGY, 2000, 18 (06) :609-613
[20]   Detection of Large Numbers of Novel Sequences in the Metatranscriptomes of Complex Marine Microbial Communities [J].
Gilbert, Jack A. ;
Field, Dawn ;
Huang, Ying ;
Edwards, Rob ;
Li, Weizhong ;
Gilna, Paul ;
Joint, Ian .
PLOS ONE, 2008, 3 (08)