Text Mining in Clinical Domain: Dealing with Noise

被引:14
作者
Hoang Nguyen [1 ]
Patrick, Jon [2 ]
机构
[1] CSIRO, Data61, 13 Garden St, Eveleigh, NSW 2015, Australia
[2] Univ Sydney, 1 Cleveland St, Sydney, NSW 2006, Australia
来源
KDD'16: PROCEEDINGS OF THE 22ND ACM SIGKDD INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA MINING | 2016年
关键词
Clinical; active learning; text classification; named-entity recognition; natural languages processing; INFORMATION;
D O I
10.1145/2939672.2939720
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Text mining in clinical domain is usually more difficult than general domains (e.g. newswire reports and scientific literature) because of the high level of noise in both the corpus and training data for machine learning (ML). A large number of unknown word, non-word and poor grammatical sentences made up the noise in the clinical corpus. Unknown words are usually complex medical vocabularies, misspellings, acronyms and abbreviations where unknown non words are generally the clinical patterns including scores and measures. This noise produces obstacles in the initial lexical processing step as well as subsequent semantic analysis. Furthermore, the labelled data used to build ML models is very costly to obtain because it requires intensive clinical knowledge from the annotators. And even created by experts, the training examples usually contain errors and inconsistencies due to the variations in human annotators' attentiveness. Clinical domain also suffers from the nature of the imbalanced data distribution problem. These kinds of noise are very popular and potentially affect the overall information extraction performance but they were not carefully investigated in most presented health informatics systems. This paper introduces a general clinical data mining architecture which is potential of addressing all of these challenges using: automatic proof-reading process, trainable finite state pattern recogniser, iterative model development and active learning. The reportability classifier based on this architecture achieved 98.25% sensitivity and 96.14% specificity on an Australian cancer registry's held-out test set and up to 92% of training data provided for supervised ML was saved by active learning.
引用
收藏
页码:549 / 558
页数:10
相关论文
共 36 条
  • [21] Collection of cancer stage data by classifying free-text medical reports
    McCowan, Iain A.
    Moore, Darren C.
    Nguyen, Anthony N.
    Bowman, Rayleen V.
    Clarke, Belinda E.
    Duhig, Edwina E.
    Fry, Mary-Jane
    [J]. JOURNAL OF THE AMERICAN MEDICAL INFORMATICS ASSOCIATION, 2007, 14 (06) : 736 - 745
  • [22] McCowan Ian, 2006, Conf Proc IEEE Eng Med Biol Soc, V2006, P5153
  • [23] Olsson F, 2009, P 13 C COMP NAT LANG, pp138
  • [24] Osugi T., 2012, P 5 IEEE INT C DAT M, P445
  • [25] Patrick J., 2011, P 25 PAC AS C LANG I, P303
  • [26] Patrick J., 2010, 2 WORKSH BUILD EV RE, P2
  • [27] Patrick J, 2011, LECT NOTES COMPUT SC, V6609, P151, DOI 10.1007/978-3-642-19437-5_12
  • [28] Schohn Greg, 2000, ICML, P839
  • [29] Settles B., 2009, 648 U WISC MAD
  • [30] Automatic structuring of radiology free-text reports
    Taira, RK
    Soderland, SG
    Jakobovits, RM
    [J]. RADIOGRAPHICS, 2001, 21 (01) : 237 - 245