Araport11: a complete reannotation of the Arabidopsis thaliana reference genome

被引:437
作者
Cheng, Chia-Yi [1 ]
Krishnakumar, Vivek [1 ]
Chan, Agnes P. [1 ]
Thibaud-Nissen, Francoise [2 ]
Schobel, Seth [1 ]
Town, Christopher D. [1 ]
机构
[1] J Craig Venter Inst, 9714 Med Ctr Dr, Rockville, MD 20850 USA
[2] NIH, Natl Ctr Biotechnol Informat, US Natl Lib Med, Bethesda, MD 20894 USA
基金
美国国家科学基金会;
关键词
Arabidopsis; annotation; transcriptome; NATURAL ANTISENSE TRANSCRIPTS; NONSENSE-MEDIATED DECAY; SMALL RNA LOCI; GENE-EXPRESSION; WIDE ANALYSIS; COMPREHENSIVE ANNOTATION; NONCODING RNAS; POLYMERASE-IV; SEQ DATA; REVEALS;
D O I
10.1111/tpj.13415
中图分类号
Q94 [植物学];
学科分类号
071001 ;
摘要
The flowering plant Arabidopsis thaliana is a dicot model organism for research in many aspects of plant biology. A comprehensive annotation of its genome paves the way for understanding the functions and activities of all types of transcripts, including mRNA, the various classes of non-coding RNA, and small RNA. The TAIR10 annotation update had a profound impact on Arabidopsis research but was released more than 5years ago. Maintaining the accuracy of the annotation continues to be a prerequisite for future progress. Using an integrative annotation pipeline, we assembled tissue-specific RNA-Seq libraries from 113 datasets and constructed 48359 transcript models of protein-coding genes in eleven tissues. In addition, we annotated various classes of non-coding RNA including microRNA, long intergenic RNA, small nucleolar RNA, natural antisense transcript, small nuclear RNA, and small RNA using published datasets and in-house analytic results. Altogether, we identified 635 novel protein-coding genes, 508 novel transcribed regions, 5178 non-coding RNAs, and 35846 small RNA loci that were formerly unannotated. Analysis of the splicing events and RNA-Seq based expression profiles revealed the landscapes of gene structures, untranslated regions, and splicing activities to be more intricate than previously appreciated. Furthermore, we present 692 uniformly expressed housekeeping genes, 43% of whose human orthologs are also housekeeping genes. This updated Arabidopsis genome annotation with a substantially increased resolution of gene models will not only further our understanding of the biological processes of this plant model but also of other species.
引用
收藏
页码:789 / 804
页数:16
相关论文
共 105 条
  • [1] A survey of the sorghum transcriptome using single-molecule long reads
    Abdel-Ghany, Salah E.
    Hamilton, Michael
    Jacobi, Jennifer L.
    Ngam, Peter
    Devitt, Nicholas
    Schilkey, Faye
    Ben-Hur, Asa
    Reddy, Anireddy S. N.
    [J]. NATURE COMMUNICATIONS, 2016, 7
  • [2] Leveraging transcript quantification for fast computation of alternative splicing profiles
    Alamancos, Gael P.
    Pages, Amadis
    Trincado, Juan L.
    Bellora, Nicolas
    Eyras, Eduardo
    [J]. RNA, 2015, 21 (09) : 1521 - 1531
  • [3] [Anonymous], 2015, 021592 BIORXIV
  • [4] ShortStack: Comprehensive annotation and quantification of small RNA genes
    Axtell, Michael J.
    [J]. RNA, 2013, 19 (06) : 740 - 751
  • [5] EXON RECOGNITION IN VERTEBRATE SPLICING
    BERGET, SM
    [J]. JOURNAL OF BIOLOGICAL CHEMISTRY, 1995, 270 (06) : 2411 - 2414
  • [6] High-resolution experimental and computational profiling of tissue-specific known and novel miRNAs in Arabidopsis
    Breakfield, Natalie W.
    Corcoran, David L.
    Petricka, Jalean J.
    Shen, Jeffrey
    Sae-Seaw, Juthamas
    Rubio-Somoza, Ignacio
    Weigel, Detlef
    Ohler, Uwe
    Benfey, Philip N.
    [J]. GENOME RESEARCH, 2012, 22 (01) : 163 - 176
  • [7] Diversity and dynamics of the Drosophila transcriptome
    Brown, James B.
    Boley, Nathan
    Eisman, Robert
    May, Gemma E.
    Stoiber, Marcus H.
    Duff, Michael O.
    Booth, Ben W.
    Wen, Jiayu
    Park, Soo
    Suzuki, Ana Maria
    Wan, Kenneth H.
    Yu, Charles
    Zhang, Dayu
    Carlson, Joseph W.
    Cherbas, Lucy
    Eads, Brian D.
    Miller, David
    Mockaitis, Keithanne
    Roberts, Johnny
    Davis, Carrie A.
    Frise, Erwin
    Hammonds, Ann S.
    Olson, Sara
    Shenker, Sol
    Sturgill, David
    Samsonova, Anastasia A.
    Weiszmann, Richard
    Robinson, Garret
    Hernandez, Juan
    Andrews, Justen
    Bickel, Peter J.
    Carninci, Piero
    Cherbas, Peter
    Gingeras, Thomas R.
    Hoskins, Roger A.
    Kaufman, Thomas C.
    Lai, Eric C.
    Oliver, Brian
    Perrimon, Norbert
    Graveley, Brenton R.
    Celniker, Susan E.
    [J]. NATURE, 2014, 512 (7515) : 393 - 399
  • [8] Improved definition of the mouse transcriptome via targeted RNA sequencing
    Bussotti, Giovanni
    Leonardi, Tommaso
    Clark, Michael B.
    Mercer, Tim R.
    Crawford, Joanna
    Malquori, Lorenzo
    Notredame, Cedric
    Dinger, Marcel E.
    Mattick, John S.
    Enright, Anton J.
    [J]. GENOME RESEARCH, 2016, 26 (05) : 705 - 716
  • [9] MAKER-P: A Tool Kit for the Rapid Creation, Management, and Quality Control of Plant Genome Annotations
    Campbell, Michael S.
    Law, MeiYee
    Holt, Carson
    Stein, Joshua C.
    Moghe, Gaurav D.
    Hufnagel, David E.
    Lei, Jikai
    Achawanantakun, Rujira
    Jiao, Dian
    Lawrence, Carolyn J.
    Ware, Doreen
    Shiu, Shin-Han
    Childs, Kevin L.
    Sun, Yanni
    Jiang, Ning
    Yandell, Mark
    [J]. PLANT PHYSIOLOGY, 2014, 164 (02) : 513 - 524
  • [10] PlantRNA, a database for tRNAs of photosynthetic eukaryotes
    Cognat, Valerie
    Pawlak, Gael
    Duchene, Anne-Marie
    Daujat, Magali
    Gigant, Anais
    Salinas, Thalia
    Michaud, Morgane
    Gutmann, Bernard
    Giege, Philippe
    Gobert, Anthony
    Marechal-Drouard, Laurence
    [J]. NUCLEIC ACIDS RESEARCH, 2013, 41 (D1) : D273 - D279