TrigNER: automatically optimized biomedical event trigger recognition on scientific documents

被引:13
作者
Campos, David [1 ]
Bui, Quoc-Chinh [2 ]
Matos, Sergio [1 ]
Oliveira, Jose Luis [1 ]
机构
[1] Univ Aveiro, IEETA DETI, P-3810 Aveiro, Portugal
[2] Erasmus MC, Dept Med Informat, Rotterdam, Netherlands
关键词
D O I
10.1186/1751-0473-9-1
中图分类号
Q [生物科学];
学科分类号
07 ; 0710 ; 09 ;
摘要
Background: Cellular events play a central role in the understanding of biological processes and functions, providing insight on both physiological and pathogenesis mechanisms. Automatic extraction of mentions of such events from the literature represents an important contribution to the progress of the biomedical domain, allowing faster updating of existing knowledge. The identification of trigger words indicating an event is a very important step in the event extraction pipeline, since the following task(s) rely on its output. This step presents various complex and unsolved challenges, namely the selection of informative features, the representation of the textual context, and the selection of a specific event type for a trigger word given this context. Results: We propose TrigNER, a machine learning-based solution for biomedical event trigger recognition, which takes advantage of Conditional Random Fields (CRFs) with a high-end feature set, including linguistic-based, orthographic, morphological, local context and dependency parsing features. Additionally, a completely configurable algorithm is used to automatically optimize the feature set and training parameters for each event type. Thus, it automatically selects the features that have a positive contribution and automatically optimizes the CRF model order, n-grams sizes, vertex information and maximum hops for dependency parsing features. The final output consists of various CRF models, each one optimized to the linguistic characteristics of each event type. Conclusions: TrigNER was tested in the BioNLP 2009 shared task corpus, achieving a total F-measure of 62.7 and outperforming existing solutions on various event trigger types, namely gene expression, transcription, protein catabolism, phosphorylation and binding. The proposed solution allows researchers to easily apply complex and optimized techniques in the recognition of biomedical event triggers, making its application a simple routine task. We believe this work is an important contribution to the biomedical text mining community, contributing to improved and faster event recognition on scientific articles, and consequent hypothesis generation and knowledge discovery. This solution is freely available as open source at http://bioinformatics.ua.pt/trigner.
引用
收藏
页数:13
相关论文
共 22 条
[1]   Gene Ontology: tool for the unification of biology [J].
Ashburner, M ;
Ball, CA ;
Blake, JA ;
Botstein, D ;
Butler, H ;
Cherry, JM ;
Davis, AP ;
Dolinski, K ;
Dwight, SS ;
Eppig, JT ;
Harris, MA ;
Hill, DP ;
Issel-Tarver, L ;
Kasarskis, A ;
Lewis, S ;
Matese, JC ;
Richardson, JE ;
Ringwald, M ;
Rubin, GM ;
Sherlock, G .
NATURE GENETICS, 2000, 25 (01) :25-29
[2]  
Bjorne J., 2009, P BIONLP 2009 WORKSH, P10
[3]   University of Turku in the BioNLP'11 Shared Task [J].
Bjorne, Jari ;
Ginter, Filip ;
Salakoski, Tapio .
BMC BIOINFORMATICS, 2012, 13 :S4
[4]   The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003 [J].
Boeckmann, B ;
Bairoch, A ;
Apweiler, R ;
Blatter, MC ;
Estreicher, A ;
Gasteiger, E ;
Martin, MJ ;
Michoud, K ;
O'Donovan, C ;
Phan, I ;
Pilbout, S ;
Schneider, M .
NUCLEIC ACIDS RESEARCH, 2003, 31 (01) :365-370
[5]   A modular framework for biomedical concept recognition [J].
Campos, David ;
Matos, Sergio ;
Oliveira, Jose Luis .
BMC BIOINFORMATICS, 2013, 14
[6]   Gimli: open source and high-performance biomedical name recognition [J].
Campos, David ;
Matos, Sergio ;
Oliveira, Jose Luis .
BMC BIOINFORMATICS, 2013, 14
[7]  
Casillas A., 2011, P BIONLP SHARED TASK, P138
[8]  
Dijkstra E.W., 1959, NUMER MATH, V1, P269, DOI DOI 10.1007/BF01386390
[9]  
Kilicoglu H, 2009, SYNTACTIC DEPENDENCY
[10]  
Kim J-D, 2011, ASS COMPUTATIONAL LI