Suitability of random forest analysis for epidemiological research: Exploring sociodemographic and lifestyle-related risk factors of overweight in a cross-sectional design

被引:23
作者
Kanerva, Noora [1 ,2 ]
Kontto, Jukka [2 ]
Erkkola, Maijaliisa [3 ]
Nevalainen, Jaakko [4 ]
Mannisto, Satu [2 ]
机构
[1] Univ Helsinki, Dept Publ Hlth, POB 20, Helsinki 00140, Finland
[2] Natl Inst Hlth & Welf, Dept Publ Hlth Solut, Helsinki, Finland
[3] Univ Helsinki, Nutr Unit, Helsinki, Finland
[4] Univ Tampere, Sch Hlth Sci, Tampere, Finland
关键词
Machine learning; mutual importance; obesity; random forest; risk factor; FOOD FREQUENCY QUESTIONNAIRE; VALIDITY; CLASSIFICATION; REGRESSION;
D O I
10.1177/1403494817736944
中图分类号
R1 [预防医学、卫生学];
学科分类号
1004 ; 120402 ;
摘要
Aims: Factors that contribute to the development of overweight are numerous and form a complex structure with many unknown interactions and associations. We aimed to explore this structure (i.e. the mutual importance or hierarchy of sociodemographic and lifestyle-related risk factors of being overweight) using a machine-learning technique called random forest (RF). The results were compared with traditional logistic regression (LR) analysis. Methods: The cross-sectional FINRISK 2007 Study included 4757 Finns (aged 25-74 years). Information on participants' lifestyle and sociodemographic characteristics were collected with questionnaires. Diet was assessed, using a validated food-frequency questionnaire. Height and weight were measured. Participants with a body mass index (BMI) 25 kg/m(2) were classified as overweight. R-statistical software was used to run RF analysis (randomForest') to derive estimates for variable importance and out-of-bag error, which were compared to a LR model. Results: In total, 704 (32%) men and 1119 (44%) women had normal BMI, whereas 1502 (69%) men and 1432 (57%) women had BMI 25. Estimated error rates for the models were similar (RF vs. LR: 42% vs. 40% for men, 38% vs. 35% for women). Both models ranked age, education and physical activity as the most important risk factors for being overweight, but RF ranked macronutrients (carbohydrates and protein) as more important compared to LR. Conclusions: RF did not demonstrate higher power in variable selection compared to LR in our study. The features of RF are more likely to appear beneficial in settings with a larger number of predictors.
引用
收藏
页码:557 / 564
页数:8
相关论文
共 30 条
[1]  
[Anonymous], 2019, R: A language for environment for statistical computing
[2]  
[Anonymous], 2008, NATL FINDIET 2007 SU
[3]  
[Anonymous], 2000, WHO TECHN REP SER
[4]   SmcHD1, containing a structural-maintenance-of-chromosomes hinge domain, has a critical role in X inactivation [J].
Blewitt, Marnie E. ;
Gendrel, Anne-Valerie ;
Pang, Zhenyi ;
Sparrow, Duncan B. ;
Whitelaw, Nadia ;
Craig, Jeffrey M. ;
Apedaile, Anwyn ;
Hilton, Douglas J. ;
Dunwoodie, Sally L. ;
Brockdorff, Neil ;
Kay, Graham F. ;
Whitelaw, Emma .
NATURE GENETICS, 2008, 40 (05) :663-669
[5]   Random forests [J].
Breiman, L .
MACHINE LEARNING, 2001, 45 (01) :5-32
[6]   The Effects of the Mediterranean Diet on Biomarkers of Vascular Wall Inflammation and Plaque Vulnerability in Subjects with High Risk for Cardiovascular Disease. A Randomized Trial [J].
Casas, Rosa ;
Sacanella, Emilio ;
Urpi-Sarda, Mireia ;
Chiva-Blanch, Gemma ;
Ros, Emilio ;
Martinez-Gonzalez, Miguel-Angel ;
Covas, Maria-Isabel ;
Ma Lamuela-Raventos, Rosa ;
Salas-Salvado, Jordi ;
Fiol, Miquel ;
Aros, Fernando ;
Estruch, Ramon .
PLOS ONE, 2014, 9 (06)
[7]   Random forests for genomic data analysis [J].
Chen, Xi ;
Ishwaran, Hemant .
GENOMICS, 2012, 99 (06) :323-329
[8]   Multicenter Comparison of Machine Learning Methods and Conventional Regression for Predicting Clinical Deterioration on the Wards [J].
Churpek, Matthew M. ;
Yuen, Trevor C. ;
Winslow, Christopher ;
Meltzer, David O. ;
Kattan, Michael W. ;
Edelson, Dana P. .
CRITICAL CARE MEDICINE, 2016, 44 (02) :368-374
[9]  
Crawford D, 2002, ASIA PAC J CLIN NUTR, V11, pS718, DOI 10.1046/j.1440-6047.11.s8.14.x
[10]   National, regional, and global trends in body-mass index since 1980: systematic analysis of health examination surveys and epidemiological studies with 960 country-years and 9.1 million participants [J].
Finucane, Mariel M. ;
Stevens, Gretchen A. ;
Cowan, Melanie J. ;
Danaei, Goodarz ;
Lin, John K. ;
Paciorek, Christopher J. ;
Singh, Gitanjali M. ;
Gutierrez, Hialy R. ;
Lu, Yuan ;
Bahalim, Adil N. ;
Farzadfar, Farshad ;
Riley, Leanne M. ;
Ezzati, Majid .
LANCET, 2011, 377 (9765) :557-567