Natural Language Processing for Improved Characterization of COVID-19 Symptoms: Observational Study of 350,000 Patients in a Large Integrated Health Care System

被引:15
作者
Malden, Deborah E. [1 ,2 ]
Tartof, Sara Y. [2 ,3 ]
Ackerson, Bradley K. [4 ]
Hong, Vennis [2 ]
Skarbinski, Jacek [5 ,6 ]
Yau, Vincent [7 ,8 ]
Qian, Lei [2 ]
Fischer, Heidi [2 ]
Shaw, Sally F. [2 ]
Caparosa, Susan [2 ]
Xie, Fagen [2 ]
机构
[1] Ctr Dis Control & Prevent, Epidem Intelligence Serv, Atlanta, GA USA
[2] Kaiser Permanente Southern Calif, Dept Res & Evaluat, 100 S Los Robles,2nd Floor, Pasadena, CA 91101 USA
[3] Kaiser Permanente Bernard J Tyson Sch Med, Pasadena, CA USA
[4] Southern Calif Permanente Med Grp, Harbor City, CA USA
[5] Kaiser Permanente Northern Calif, Permanente Med Grp, Oakland, CA USA
[6] Kaiser Permanente Northern Calif, Div Res, Oakland, CA USA
[7] Genentech Inc, San Francisco, CA USA
[8] Roche Grp, San Francisco, CA USA
关键词
natural language processing; NLP; COVID-19; symptoms; disease characterization; artificial intelligence; application; data; cough; fever; headache; surveillance; STATES;
D O I
10.2196/41529
中图分类号
R1 [预防医学、卫生学];
学科分类号
1004 ; 120402 ;
摘要
Background: Natural language processing (NLP) of unstructured text from electronic medical records (EMR) can improve the characterization of COVID-19 signs and symptoms, but large-scale studies demonstrating the real-world application and validation of NLP for this purpose are limited.Objective: The aim of this paper is to assess the contribution of NLP when identifying COVID-19 signs and symptoms from EMR.Methods: This study was conducted in Kaiser Permanente Southern California, a large integrated health care system using data from all patients with positive SARS-CoV-2 laboratory tests from March 2020 to May 2021. An NLP algorithm was developed to extract free text from EMR on 12 established signs and symptoms of COVID-19, including fever, cough, headache, fatigue, dyspnea, chills, sore throat, myalgia, anosmia, diarrhea, vomiting or nausea, and abdominal pain. The proportion of patients reporting each symptom and the corresponding onset dates were described before and after supplementing structured EMR data with NLP-extracted signs and symptoms. A random sample of 100 chart-reviewed and adjudicated SARS-CoV-2-positive cases were used to validate the algorithm performance.Results: A total of 359,938 patients (mean age 40.4 [SD 19.2] years; 191,630/359,938, 53% female) with confirmed SARS-CoV-2 infection were identified over the study period. The most common signs and symptoms identified through NLP-supplemented analyses were cough (220,631/359,938, 61%), fever (185,618/359,938, 52%), myalgia (153,042/359,938, 43%), and headache (144,705/359,938, 40%). The NLP algorithm identified an additional 55,568 (15%) symptomatic cases that were previously defined as asymptomatic using structured data alone. The proportion of additional cases with each selected symptom identified in NLP-supplemented analysis varied across the selected symptoms, from 29% (63,742/220,631) of all records for cough to 64% (38,884/60,865) of all records with nausea or vomiting. Of the 295,305 symptomatic patients, the median time from symptom onset to testing was 3 days using structured data alone, whereas the NLP algorithm identified signs or symptoms approximately 1 day earlier. When validated against chart-reviewed cases, the NLP algorithm successfully identified signs and symptoms with consistently high sensitivity (ranging from 87% to 100%) and specificity (94% to 100%).Conclusions: These findings demonstrate that NLP can identify and characterize a broad set of COVID-19 signs and symptoms from unstructured EMR data with enhanced detail and timeliness compared with structured data alone.
引用
收藏
页数:13
相关论文
共 37 条
[1]  
Alhussayni KH, 2021, PERIODI ENG NAT SCI, V9, P667, DOI 10.21533/pen.v9i2.1862
[2]   Population-scale longitudinal mapping of COVID-19 symptoms, behaviour and testing [J].
Allen, William E. ;
Altae-Tran, Han ;
Briggs, James ;
Jin, Xin ;
McGee, Glen ;
Shi, Andy ;
Raghavan, Rumya ;
Kamariza, Mireille ;
Nova, Nicole ;
Pereta, Albert ;
Danford, Chris ;
Kamel, Amine ;
Gothe, Patrik ;
Milam, Evrhet ;
Aurambault, Jean ;
Primke, Thorben ;
Li, Weijie ;
Inkenbrandt, Josh ;
Tuan Huynh ;
Chen, Evan ;
Lee, Christina ;
Croatto, Michael ;
Bentley, Helen ;
Lu, Wendy ;
Murray, Robert ;
Travassos, Mark ;
Coull, Brent A. ;
Openshaw, John ;
Greene, Casey S. ;
Shalem, Ophir ;
King, Gary ;
Probasco, Ryan ;
Cheng, David R. ;
Silbermann, Ben ;
Zhang, Feng ;
Lin, Xihong .
NATURE HUMAN BEHAVIOUR, 2020, 4 (09) :972-+
[3]   Evidence of Gender Differences in the Diagnosis and Management of Coronavirus Disease 2019 Patients: An Analysis of Electronic Health Records Using Natural Language Processing and Machine Learning [J].
Ancochea, Julio ;
Izquierdo, Jose L. ;
Soriano, Joan B. .
JOURNAL OF WOMENS HEALTH, 2021, 30 (03) :393-404
[4]  
[Anonymous], 2022, SYMPT COVID 19
[5]  
[Anonymous], WHO COVID-19 dashboard, COVID-19 vaccination, World data
[6]   Web Search Engine Misinformation Notifier Extension (SEMiNExt): A Machine Learning Based Approach during COVID-19 Pandemic [J].
Bin Shams, Abdullah ;
Hoque Apu, Ehsanul ;
Rahman, Ashiqur ;
Sarker Raihan, Md. Mohsin ;
Siddika, Nazeeba ;
Bin Preo, Rahat ;
Hussein, Molla Rashied ;
Mostari, Shabnam ;
Kabir, Russell .
HEALTHCARE, 2021, 9 (02)
[7]   Development and validation of an automated emergency department-based syndromic surveillance system to enhance public health surveillance in Yukon: a lower-resourced and remote setting [J].
Bouchouar, Etran ;
Hetman, Benjamin M. ;
Hanley, Brendan .
BMC PUBLIC HEALTH, 2021, 21 (01)
[8]   Symptom Profiles of a Convenience Sample of Patients with COVID-19-United States, January-April 2020 [J].
Burke, Rachel M. ;
Killerby, Marie E. ;
Newton, Suzanne ;
Ashworth, Candace E. ;
Berns, Abby L. ;
Brennan, Skyler ;
Bressler, Jonathan M. ;
Bye, Erica ;
Crawford, Richard ;
Morano, Laurel Harduar ;
Lewis, Nathaniel M. ;
Markus, Tiffanie M. ;
Read, Jennifer S. ;
Rissman, Tamara ;
Taylor, Joanne ;
Tate, Jacqueline E. ;
Midgley, Claire M. .
MMWR-MORBIDITY AND MORTALITY WEEKLY REPORT, 2020, 69 (28) :904-908
[9]   SARS-CoV-2, SARS-CoV, and MERS-CoV viral load dynamics, duration of viral shedding, and infectiousness: a systematic review and meta-analysis [J].
Cevik, Muge ;
Tate, Matthew ;
Lloyd, Ollie ;
Maraolo, Alberto Enrico ;
Schafers, Jenna ;
Ho, Antonia .
LANCET MICROBE, 2021, 2 (01) :E13-E22
[10]   New options for national population surveys: The implications of internet and smartphone coverage [J].
Couper, Mick P. ;
Gremel, Garret ;
Axinn, William ;
Guyer, Heidi ;
Wagner, James ;
West, Brady T. .
SOCIAL SCIENCE RESEARCH, 2018, 73 :221-235