共 50 条
Discovery of Novel Sequences in 1,000 Swedish Genomes
被引:19
作者:
Eisfeldt, Jesper
[1
,2
,3
]
Martensson, Gustaf
[4
]
Ameur, Adam
[5
]
Nilsson, Daniel
[1
,2
,3
]
Lindstrand, Anna
[1
,3
]
机构:
[1] Karolinska Inst, Ctr Mol Med, Dept Mol Med & Surg, Stockholm, Sweden
[2] Karolinska Inst, Sci Life Lab, Sci Pk, Solna, Sweden
[3] Karolinska Univ Hosp, Dept Clin Genet, Stockholm, Sweden
[4] KTH Royal Inst Technol, Sch Engn Sci Chem Biotechnol & Hlth, Sci Life Lab, Div Nanobiotechnol,Dept Prot Sci, Stockholm, Sweden
[5] Uppsala Univ, Sci Life Lab, Dept Immunol Genet & Pathol, Uppsala, Sweden
基金:
瑞典研究理事会;
关键词:
population genomics;
novel sequences;
de novo assembly;
ancestral deletion;
GENETIC-VARIATION;
PROTEIN;
DIVERSITY;
ALIGNMENT;
RESOURCE;
D O I:
10.1093/molbev/msz176
中图分类号:
Q5 [生物化学];
Q7 [分子生物学];
学科分类号:
071010 ;
081704 ;
摘要:
Novel sequences (NSs), not present in the human reference genome, are abundant and remain largely unexplored. Here, we utilize de novo assembly to study NS in 1,000 Swedish individuals first sequenced as part of the SweGen project revealing a total of 46 Mb in 61,044 distinct contigs of sequences not present in GRCh38. The contigs were aligned to recently published catalogs of Icelandic and Pan-African NSs, as well as the chimpanzee genome, revealing a great diversity of shared sequences. Analyzing the positioning of NS across the chimpanzee genome, we find that 2,807 NS align confidently within 143 chimpanzee orthologs of human genes. Aligning the whole genome sequencing data to the chimpanzee genome, we discover ancestral NS common throughout the Swedish population. The NSs were searched for repeats and repeat elements: revealing a majority of repetitive sequence (56%), and enrichment of simple repeats (28%) and satellites (15%). Lastly, we align the unmappable reads of a subset of the thousand genomes data to our collection of NS, as well as the previously published Pan-African NS: revealing that both the Swedish and Pan-African NS are widespread, and that the Swedish NSs are largely a subset of the Pan-African NS. Overall, these results highlight the importance of creating a more diverse reference genome and illustrate that significant amounts of the NS may be of ancestral origin.
引用
收藏
页码:18 / 30
页数:13
相关论文