Rule-based information extraction from patients' clinical data

被引:88
作者
Mykowiecka, Agnieszka [1 ]
Marciniak, Malgorzata [1 ]
Kupsc, Anna [1 ]
机构
[1] Inst Comp Sci PAS, PL-01237 Warsaw, Poland
关键词
Rule-based information extraction; Polish clinical data; Linguistic analysis;
D O I
10.1016/j.jbi.2009.07.007
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
The paper describes a rule-based information extraction (IE) system developed for Polish medical texts. We present two applications designed to select data from medical documentation in Polish: mammography reports and hospital records of diabetic patients. First, we have designed a special ontology that subsequently had its concepts translated into two separate models, represented as typed feature structure (TFS) hierarchies, complying with the format required by the IE platform we adopted. Then, we used dedicated IE grammars to process documents and fill in templates provided by the models. In particular, in the grammars, we addressed such linguistic issues as: ambiguous keywords, negation, coordination or anaphoric expressions. Resolving some of these problems has been deferred to a post-processing phase where the extracted information is further grouped and structured into more complex templates. To this end, we defined special heuristic algorithms on the basis of sample data. The evaluation of the implemented procedures shows their usability for clinical data extraction tasks. For most of the evaluated templates, precision and recall well above 80% were obtained. (C) 2009 Elsevier Inc. All rights reserved.
引用
收藏
页码:923 / 936
页数:14
相关论文
共 43 条
[1]  
Agirre Eneko, 2007, Word Sense Disambiguation : Algorithms and Applications, V1st
[2]  
[Anonymous], SHALLOW PROCESSING U
[3]  
[Anonymous], 1992, The logic of typed feature structures
[4]  
BONTCHEVA K, 2002, P TRAIT AUT LANG NAT
[5]  
Boytcheva S., 2005, International Workshop Language and Speech Infrastructure for Information Access in the Balkan Countries, Borovets, Bulgaria, P1
[6]  
BUNESCU R, 2003, P ICML 2003 WORKSH M, P46
[7]  
BURNSIDE E, 2000, P AM MED INF ASS S, P16
[8]  
BURNSIDE E, 2000, CARS 2000 INT C COMP
[9]  
BUYKO E, 2007, P 10 C PAC ASS COMP
[10]  
CHAPMAN VVW, 2001, J BIOMED INFORM, V34, P204