A shortest path-based approach for copy number variation detection from next-generation sequencing data

被引:4
作者
Liu, Guojun [1 ]
Yang, Hongzhi [2 ]
Yuan, Xiguo [3 ]
机构
[1] Xian Univ Finance & Econ, Sch Stat, Xian, Peoples R China
[2] Xidian Grp Hosp, Med Imaging Ctr, Xian, Peoples R China
[3] Xidian Univ, Hangzhou Inst Technol, Hangzhou, Peoples R China
关键词
copy number variation; next-generation sequencing data; k nearest neighbors; shortest path; read depth; DUPLICATIONS; ALGORITHM; DELETIONS;
D O I
10.3389/fgene.2022.1084974
中图分类号
Q3 [遗传学];
学科分类号
071007 ; 090102 ;
摘要
Copy number variation (CNV) is one of the main structural variations in the human genome and accounts for a considerable proportion of variations. As CNVs can directly or indirectly cause cancer, mental illness, and genetic disease in humans, their effective detection in humans is of great interest in the fields of oncogene discovery, clinical decision-making, bioinformatics, and drug discovery. The advent of next-generation sequencing data makes CNV detection possible, and a large number of CNV detection tools are based on next-generation sequencing data. Due to the complexity (e.g., bias, noise, alignment errors) of next-generation sequencing data and CNV structures, the accuracy of existing methods in detecting CNVs remains low. In this work, we design a new CNV detection approach, called shortest path-based Copy number variation (SPCNV), to improve the detection accuracy of CNVs. SPCNV calculates the k nearest neighbors of each read depth and defines the shortest path, shortest path relation, and shortest path cost sets based on which further calculates the mean shortest path cost of each read depth and its k nearest neighbors. We utilize the ratio between the mean shortest path cost for each read depth and the mean of the mean shortest path cost of its k nearest neighbors to construct a relative shortest path score formula that is able to determine a score for each read depth. Based on the score profile, a boxplot is then applied to predict CNVs. The performance of the proposed method is verified by simulation data experiments and compared against several popular methods of the same type. Experimental results show that the proposed method achieves the best balance between recall and precision in each set of simulated samples. To further verify the performance of the proposed method in real application scenarios, we then select real sample data from the 1,000 Genomes Project to conduct experiments. The proposed method achieves the best F1-scores in almost all samples. Therefore, the proposed method can be used as a more reliable tool for the routine detection of CNVs.
引用
收藏
页数:10
相关论文
共 39 条
  • [1] CNVnator: An approach to discover, genotype, and characterize typical and atypical CNVs from family and population genome sequencing
    Abyzov, Alexej
    Urban, Alexander E.
    Snyder, Michael
    Gerstein, Mark
    [J]. GENOME RESEARCH, 2011, 21 (06) : 974 - 984
  • [2] The landscape of somatic copy-number alteration across human cancers
    Beroukhim, Rameen
    Mermel, Craig H.
    Porter, Dale
    Wei, Guo
    Raychaudhuri, Soumya
    Donovan, Jerry
    Barretina, Jordi
    Boehm, Jesse S.
    Dobson, Jennifer
    Urashima, Mitsuyoshi
    Mc Henry, Kevin T.
    Pinchback, Reid M.
    Ligon, Azra H.
    Cho, Yoon-Jae
    Haery, Leila
    Greulich, Heidi
    Reich, Michael
    Winckler, Wendy
    Lawrence, Michael S.
    Weir, Barbara A.
    Tanaka, Kumiko E.
    Chiang, Derek Y.
    Bass, Adam J.
    Loo, Alice
    Hoffman, Carter
    Prensner, John
    Liefeld, Ted
    Gao, Qing
    Yecies, Derek
    Signoretti, Sabina
    Maher, Elizabeth
    Kaye, Frederic J.
    Sasaki, Hidefumi
    Tepper, Joel E.
    Fletcher, Jonathan A.
    Tabernero, Josep
    Baselga, Jose
    Tsao, Ming-Sound
    Demichelis, Francesca
    Rubin, Mark A.
    Janne, Pasi A.
    Daly, Mark J.
    Nucera, Carmelo
    Levine, Ross L.
    Ebert, Benjamin L.
    Gabriel, Stacey
    Rustgi, Anil K.
    Antonescu, Cristina R.
    Ladanyi, Marc
    Letai, Anthony
    [J]. NATURE, 2010, 463 (7283) : 899 - 905
  • [3] Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing data
    Boeva, Valentina
    Popova, Tatiana
    Bleakley, Kevin
    Chiche, Pierre
    Cappo, Julie
    Schleiermacher, Gudrun
    Janoueix-Lerosey, Isabelle
    Delattre, Olivier
    Barillot, Emmanuel
    [J]. BIOINFORMATICS, 2012, 28 (03) : 423 - 425
  • [4] LOF: Identifying density-based local outliers
    Breunig, MM
    Kriegel, HP
    Ng, RT
    Sander, J
    [J]. SIGMOD RECORD, 2000, 29 (02) : 93 - 104
  • [5] SeqCNV: a novel method for identification of copy number variations in targeted next-generation sequencing data
    Chen, Yong
    Zhao, Li
    Wang, Yi
    Cao, Ming
    Gelowani, Violet
    Xu, Mingchu
    Agrawal, Smriti A.
    Li, Yumei
    Daiger, Stephen P.
    Gibbs, Richard
    Wang, Fei
    Chen, Rui
    [J]. BMC BIOINFORMATICS, 2017, 18
  • [6] A Direct Algorithm for 1-D Total Variation Denoising
    Condat, Laurent
    [J]. IEEE SIGNAL PROCESSING LETTERS, 2013, 20 (11) : 1054 - 1057
  • [7] iCopyDAV: Integrated platform for copy number variations-Detection, annotation and visualization
    Dharanipragada, Prashanthi
    Vogeti, Sriharsha
    Parekh, Nita
    [J]. PLOS ONE, 2018, 13 (04):
  • [8] CNV-TV: A robust method to discover copy number variation from short sequencing reads
    Duan, Junbo
    Zhang, Ji-Gang
    Deng, Hong-Wen
    Wang, Yu-Ping
    [J]. BMC BIOINFORMATICS, 2013, 14
  • [9] Copy number variation: New insights in genome diversity
    Freeman, Jennifer L.
    Perry, George H.
    Feuk, Lars
    Redon, Richard
    McCarroll, Steven A.
    Altshuler, David M.
    Aburatani, Hiroyuki
    Jones, Keith W.
    Tyler-Smith, Chris
    Hurles, Matthew E.
    Carter, Nigel P.
    Scherer, Stephen W.
    Lee, Charles
    [J]. GENOME RESEARCH, 2006, 16 (08) : 949 - 961
  • [10] Fridley Brooke L., 2012, Frontiers in Genetics, V3, P142, DOI 10.3389/fgene.2012.00142