Learning to classify species with barcodes

被引:68
作者
Bertolazzi, Paola [1 ]
Felici, Giovanni [1 ]
Weitschek, Emanuel [1 ]
机构
[1] CNR, Ist Analisi Sistemi & Informat Antonio Ruberti, I-00185 Rome, Italy
来源
BMC BIOINFORMATICS | 2009年 / 10卷
关键词
MITOCHONDRIAL COI; DNA; GENES;
D O I
10.1186/1471-2105-10-S14-S7
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Background: According to many field experts, specimens classification based on morphological keys needs to be supported with automated techniques based on the analysis of DNA fragments. The most successful results in this area are those obtained from a particular fragment of mitochondrial DNA, the gene cytochrome c oxidase I (COI) (the "barcode"). Since 2004 the Consortium for the Barcode of Life (CBOL) promotes the collection of barcode specimens and the development of methods to analyze the barcode for several tasks, among which the identification of rules to correctly classify an individual into its species by reading its barcode. Results: We adopt a Logic Mining method based on two optimization models and present the results obtained on two datasets where a number of COI fragments are used to describe the individuals that belong to different species. The method proposed exhibits high correct recognition rates on a training-testing split of the available data using a small proportion of the information available (e. g., correct recognition approx. 97% when only 20 sites of the 648 available are used). The method is able to provide compact formulas on the values (A, C, G, T) at the selected sites that synthesize the characteristic of each species, a relevant information for taxonomists. Conclusion: We have presented a Logic Mining technique designed to analyze barcode data and to provide detailed output of interest to the taxonomists and the barcode community represented in the CBOL Consortium. The method has proven to be effective, efficient and precise.
引用
收藏
页数:12
相关论文
共 31 条
  • [1] A step toward barcoding life: A model-based, decision-theoretic method to assign genes to preexisting species groups
    Abdo, Zaid
    Golding, G. Brian
    [J]. SYSTEMATIC BIOLOGY, 2007, 56 (01) : 44 - 56
  • [2] Logic classification and feature selection for biomedical data
    Bertolazzi, P.
    Felici, G.
    Festa, P.
    Lancia, G.
    [J]. COMPUTERS & MATHEMATICS WITH APPLICATIONS, 2008, 55 (05) : 889 - 899
  • [3] Mitochondrial COI and II provide useful markers for Wiseana (Lepidoptera: Hepialidae) species identification
    Brown, B
    Emberson, RM
    Paterson, AM
    [J]. BULLETIN OF ENTOMOLOGICAL RESEARCH, 1999, 89 (04) : 287 - 293
  • [4] Taxonomic and systematic assessment of planktonic copepods using mitochondrial COI sequence variation and competitive, species-specific PCR
    Bucklin, A
    Guarnieri, M
    Hill, RS
    Bentley, AM
    Kaartvedt, S
    [J]. HYDROBIOLOGIA, 1999, 401 (0) : 239 - 254
  • [5] Buschmann F., 2007, PATTERN ORIENTED SOF, V4
  • [6] CHIAJUNG C, 2006, BIOINFORMATICS, V22, P685
  • [7] FELICI G, 2006, DATA MINING KNOWLEDG
  • [8] FELICI G, 2002, INFORMS J COMPUTING, V14
  • [9] FELICI G, 2005, ENCY DATA WAREHOUSIN, P693
  • [10] Garey M. R., 1979, Computers and intractability. A guide to the theory of NP-completeness