ICD-10 Coding of Spanish Electronic Discharge Summaries: An Extreme Classification Problem

被引:19
作者
Almagro, Mario [1 ]
Martinez Unanue, Raquel [1 ]
Fresno, Victor [1 ]
Montalvo, Soto [2 ]
机构
[1] Univ Nacl Educ Distancia, Dept Comp Languages & Syst, Madrid 28040, Spain
[2] King Juan Carlos Univ URJC, Dept Comp Sci, Madrid 28933, Spain
关键词
Encoding; Training; Hospitals; Diseases; Proposals; Licenses; Task analysis; Extreme classification; XMTC; ICD-10; coding; text mining;
D O I
10.1109/ACCESS.2020.2997241
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Objective: Medical coding is used to identify and standardize clinical concepts in the records collected from healthcare services. The tenth revision of the International Classification of Diseases (ICD-10) is the most widely-used coding with more than 11,000 different diagnoses, affecting research, reporting, and funding. Unfortunately, ICD-10 code sets tend to follow biased, unbalanced, and scattered distributions. These distribution attributes, along with high lexical variability, severely restrict performance when coded clinical records are used to infer code sets in uncoded records. To improve that inference, we explore a combination of example-based methods optimized to capture codes with different appearance frequencies in data sets. Materials and Methods: The proposed exploration has been carried out on Spanish hospital discharge reports coded by experts, excluding all sentences without any biomedical concept. Representations based on semantic and lexical features are explored, using both global and label-specific attributes. In turn, algorithms based on binary outputs, groups of subsets and extreme classification are compared. Lists of codes together with their confidence values (certainty probabilities) are suggested by each method. Results: Diverse spectral behaviors are shown for each method. Binary classifiers seem to maximize the capture of more popular codes, while extreme classifiers promote infrequent ones. In order to exploit such differences, ensemble approaches are proposed by weighting every output code according to the method, confidence value and appearance frequency. The rule-based combination reaches a 46% Precision at 10 (P@10), which means a 15% improvement over the best individual proposal. Conclusion: Assembling methods based on weighting each code according to training frequency and performance can achieve better overall Precision scores on extreme distributions, such as ICD-10 coding.
引用
收藏
页码:100073 / 100083
页数:11
相关论文
共 39 条
[1]   A cross-lingual approach to automatic ICD-10 coding of death certificates by exploring machine translation [J].
Almagro, Mario ;
Martinez, Raquel ;
Montalvo, Soto ;
Fresno, Victor .
JOURNAL OF BIOMEDICAL INFORMATICS, 2019, 94
[2]   Preliminary Study of the Automatic Annotation of Hospital Discharge Report with ICD-10 codes [J].
Almagro, Mario ;
Martinez, Raquel ;
Fresno, Victor ;
Montalvo, Soto .
PROCESAMIENTO DEL LENGUAJE NATURAL, 2018, (60) :45-52
[3]  
[Anonymous], 2013, P 26 INT C NEURAL IN
[4]  
[Anonymous], 2014, P 8 INT WORKSH SEM E
[5]  
[Anonymous], 2013, ICML
[6]  
[Anonymous], 2012, ADV NEURAL INFORM PR, DOI DOI 10.1055/S-0031-1291042
[7]   CodeMagic: Semi-Automatic Assignment of ICD-10-AM Codes to Patient Records [J].
Arifoglu, Damla ;
Deniz, Onur ;
Alecakir, Kemal ;
Yondem, Meltem .
INFORMATION SCIENCES AND SYSTEMS 2014, 2014, :259-268
[8]  
Atutxa A., 2018, P CEUR WORKSH
[9]  
Baker YS, 2013, 2013 IEEE INTERNATIONAL CONFERENCE ON INTELLIGENCE AND SECURITY INFORMATICS: BIG DATA, EMERGENT THREATS, AND DECISION-MAKING IN SECURITY INFORMATICS, P10, DOI 10.1109/ISI.2013.6578776
[10]  
Balasubramanian K., 2012, ARXIV12066479