Combining Search, Social Media, and Traditional Data Sources to Improve Influenza Surveillance

被引:260
作者
Santillana, Mauricio [1 ,2 ,3 ]
Nguyen, Andre T. [1 ]
Dredze, Mark [4 ]
Paul, Michael J. [5 ]
Nsoesie, Elaine O. [6 ,7 ]
Brownstein, John S. [2 ,3 ]
机构
[1] Harvard Univ, Sch Engn & Appl Sci, Cambridge, MA 02138 USA
[2] Boston Childrens Hosp, Informat Program, Boston, MA USA
[3] Harvard Univ, Sch Med, Boston, MA USA
[4] Johns Hopkins Univ, Dept Comp Sci, Baltimore, MD 21218 USA
[5] Univ Colorado, Dept Informat Sci, Boulder, CO 80309 USA
[6] Univ Washington, Dept Global Hlth, Seattle, WA 98195 USA
[7] Inst Hlth Metr & Evaluat, Seattle, WA USA
基金
美国国家卫生研究院;
关键词
QUERY DATA; FLU; EPIDEMICS;
D O I
10.1371/journal.pcbi.1004513
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
We present a machine learning-based methodology capable of providing real-time ("now-cast") and forecast estimates of influenza activity in the US by leveraging data from multiple data sources including: Google searches, Twitter microblogs, nearly real-time hospital visit records, and data from a participatory surveillance system. Our main contribution consists of combining multiple influenza-like illnesses (ILI) activity estimates, generated independently with each data source, into a single prediction of ILI utilizing machine learning ensemble approaches. Our methodology exploits the information in each data source and produces accurate weekly ILI predictions for up to four weeks ahead of the release of CDC's ILI reports. We evaluate the predictive ability of our ensemble approach during the 2013-2014 (retrospective) and 2014-2015 (live) flu seasons for each of the four weekly time horizons. Our ensemble approach demonstrates several advantages: (1) our ensemble method's predictions outperform every prediction using each data source independently, (2) our methodology can produce predictions one week ahead of GFT's real-time estimates with comparable accuracy, and (3) our two and three week forecast estimates have comparable accuracy to real-time predictions using an autoregressive model. Moreover, our results show that considerable insight is gained from incorporating disparate data streams, in the form of social media and crowd sourced data, into influenza predictions in all time horizons.
引用
收藏
页数:15
相关论文
共 42 条
[21]   Flu Gone Viral: Syndromic Surveillance of Flu on Twitter using Temporal Topic Models [J].
Chen, Liangzhe ;
Hossain, K. S. M. Tozammel ;
Butler, Patrick ;
Ramakrishnan, Naren ;
Prakash, B. Aditya .
2014 IEEE INTERNATIONAL CONFERENCE ON DATA MINING (ICDM), 2014, :755-760
[22]   IMPROVING THE EVIDENCE BASE FOR DECISION MAKING DURING A PANDEMIC: THE EXAMPLE OF 2009 INFLUENZA A/H1N1 [J].
Lipsitch, Marc ;
Finelli, Lyn ;
Heffernan, Richard T. ;
Leung, Gabriel M. ;
Redd, Stephen C. .
BIOSECURITY AND BIOTERRORISM-BIODEFENSE STRATEGY PRACTICE AND SCIENCE, 2011, 9 (02) :89-115
[23]   A New Approach to Monitoring Dengue Activity [J].
Madoff, Lawrence C. ;
Fisman, David N. ;
Kass-Hout, Taha .
PLOS NEGLECTED TROPICAL DISEASES, 2011, 5 (05)
[24]   Wikipedia Usage Estimates Prevalence of Influenza-Like Illness in the United States in Near Real-Time [J].
McIver, David J. ;
Brownstein, John S. .
PLOS COMPUTATIONAL BIOLOGY, 2014, 10 (04)
[25]   A Case Study of the New York City 2012-2013 Influenza Season With Daily Geocoded Twitter Data From Temporal and Spatiotemporal Perspectives [J].
Nagar, Ruchit ;
Yuan, Qingyu ;
Freifeld, Clark C. ;
Santillana, Mauricio ;
Nojima, Aaron ;
Chunara, Rumi ;
Brownstein, John S. .
JOURNAL OF MEDICAL INTERNET RESEARCH, 2014, 16 (10) :260-274
[26]   Using search queries for malaria surveillance, Thailand [J].
Ocampo, Alex J. ;
Chunara, Rumi ;
Brownstein, John S. .
MALARIA JOURNAL, 2013, 12
[27]   Monitoring the impact of influenza by age: Emergency department fever and respiratory complaint surveillance in New York City [J].
Olson, Donald R. ;
Heffernan, Richard T. ;
Paladini, Marc ;
Konty, Kevin ;
Weiss, Don ;
Mostashari, Farzad .
PLOS MEDICINE, 2007, 4 (08) :1349-1361
[28]   Reassessing Google Flu Trends Data for Detection of Seasonal and Pandemic Influenza: A Comparative Epidemiological Study at Three Geographic Scales [J].
Olson, Donald R. ;
Konty, Kevin J. ;
Paladini, Marc ;
Viboud, Cecile ;
Simonsen, Lone .
PLOS COMPUTATIONAL BIOLOGY, 2013, 9 (10)
[29]   Web-based participatory surveillance of infectious diseases: the Influenzanet participatory surveillance experience [J].
Paolotti, D. ;
Carnahan, A. ;
Colizza, V. ;
Eames, K. ;
Edmunds, J. ;
Gomes, G. ;
Koppeschaar, C. ;
Rehn, M. ;
Smallenburg, R. ;
Turbelin, C. ;
Van Noort, S. ;
Vespignani, A. .
CLINICAL MICROBIOLOGY AND INFECTION, 2014, 20 (01) :17-21
[30]  
Paul M. J, 2014, PLOS CURRENTS, V6