Cardiovascular disease incidence prediction by machine learning and statistical techniques: a 16-year cohort study from eastern Mediterranean region

被引:4
作者
Mehrabani-Zeinabad, Kamran [1 ]
Feizi, Awat [2 ,3 ]
Sadeghi, Masoumeh [3 ]
Roohafza, Hamidreza [1 ]
Talaei, Mohammad [4 ]
Sarrafzadegan, Nizal [1 ,5 ]
机构
[1] Isfahan Univ Med Sci, Cardiovasc Res Inst, Cardiovasc Res Ctr, Esfahan, Iran
[2] Isfahan Univ Med Sci, Sch Hlth, Biostat & Epidemiol Dept, Esfahan, Iran
[3] Isfahan Univ Med Sci, Cardiovasc Res Inst, Cardiac Rehabil Res Ctr, Esfahan, Iran
[4] Queen Mary Univ London, Wolfson Inst Populat Hlth, Barts & London Sch Med & Dent, London, England
[5] Univ British Columbia, Fac Med, Sch Populat & Publ Hlth, Vancouver, BC, Canada
关键词
Cardiovascular; Machine learning; Statistical models; Cohort study; Eastern Mediterranean region; Feature selection; Missing values; RISK; REGRESSION; PREVENTION;
D O I
10.1186/s12911-023-02169-5
中图分类号
R-058 [];
学科分类号
摘要
BackgroundCardiovascular diseases (CVD) are the predominant cause of early death worldwide. Identification of people with a high risk of being affected by CVD is consequential in CVD prevention. This study adopts Machine Learning (ML) and statistical techniques to develop classification models for predicting the future occurrence of CVD events in a large sample of Iranians.MethodsWe used multiple prediction models and ML techniques with different abilities to analyze the large dataset of 5432 healthy people at the beginning of entrance into the Isfahan Cohort Study (ICS) (1990-2017). Bayesian additive regression trees enhanced with "missingness incorporated in attributes" (BARTm) was run on the dataset with 515 variables (336 variables without and the remaining with up to 90% missing values). In the other used classification algorithms, variables with more than 10% missing values were excluded, and MissForest imputes the missing values of the remaining 49 variables. We used Recursive Feature Elimination (RFE) to select the most contributing variables. Random oversampling technique, recommended cut-point by precision-recall curve, and relevant evaluation metrics were used for handling unbalancing in the binary response variable.ResultsThis study revealed that age, systolic blood pressure, fasting blood sugar, two-hour postprandial glucose, diabetes mellitus, history of heart disease, history of high blood pressure, and history of diabetes are the most contributing factors for predicting CVD incidence in the future. The main differences between the results of classification algorithms are due to the trade-off between sensitivity and specificity. Quadratic Discriminant Analysis (QDA) algorithm presents the highest accuracy (75.50 +/- 0.08) but the minimum sensitivity (49.84 +/- 0.25); In contrast, decision trees provide the lowest accuracy (51.95 +/- 0.69) but the top sensitivity (82.52 +/- 1.22). BARTm.90% resulted in 69.48 +/- 0.28 accuracy and 54.00 +/- 1.66 sensitivity without any preprocessing step.ConclusionsThis study confirmed that building a prediction model for CVD in each region is valuable for screening and primary prevention strategies in that specific region. Also, results showed that using conventional statistical models alongside ML algorithms makes it possible to take advantage of both techniques. Generally, QDA can accurately predict the future occurrence of CVD events with a fast (inference speed) and stable (confidence values) procedure. The combined ML and statistical algorithm of BARTm provide a flexible approach without any need for technical knowledge about assumptions and preprocessing steps of the prediction procedure.
引用
收藏
页数:12
相关论文
共 64 条
  • [1] Alaa AM, 2018, PR MACH LEARN RES, V80
  • [2] Cardiovascular disease risk prediction using automated machine learning: A prospective study of 423,604 UK Biobank participants
    Alaa, Ahmed M.
    Bolton, Thomas
    Di Angelantonio, Emanuele
    Rudd, James H. F.
    van der Schaar, Mihaela
    [J]. PLOS ONE, 2019, 14 (05):
  • [3] Reviewing the use and quality of machine learning in developing clinical prediction models for cardiovascular disease
    Allan, Simon
    Olaiya, Raphael
    Burhan, Rasan
    [J]. POSTGRADUATE MEDICAL JOURNAL, 2022, 98 (1161) : 551 - 558
  • [4] American Diabetes Association, 2022, Clin Diabetes, V40, P10, DOI 10.2337/cd22-as01
  • [5] 70-year legacy of the Framingham Heart Study
    Andersson, Charlotte
    Johnson, Andrew D.
    Benjamin, Emelia J.
    Levy, Daniel
    Vasan, Ramachandran S.
    [J]. NATURE REVIEWS CARDIOLOGY, 2019, 16 (11) : 687 - 698
  • [6] [Anonymous], 2009, AGE
  • [7] [Anonymous], NUMB ART INT AI EXP
  • [8] Arnett DK, 2019, CIRCULATION, V140, pE563, DOI [10.1161/CIR.0000000000000677, 10.1161/CIR.0000000000000678, 10.1016/j.jacc.2019.03.009, 10.1016/j.jacc.2019.03.010]
  • [9] Atkinson Beth, 2023, CRAN
  • [10] Global burden of CVD: focus on secondary prevention of cardiovascular disease
    Bansilal, Sameer
    Castellano, Jose M.
    Fuster, Valentin
    [J]. INTERNATIONAL JOURNAL OF CARDIOLOGY, 2015, 201 : S1 - S7