A Robust Machine Learning Framework for Diabetes Prediction

被引:0
作者
Olisah, Chollette [1 ]
Adeleye, Oluwaseun [2 ]
Smith, Lyndon [1 ]
Smith, Melvyn [1 ]
机构
[1] Univ West England, Ctr Machine Vis, Bristol Robot Lab, Bristol, Avon, England
[2] Baze Univ, Dept Comp Sci, Abuja, Nigeria
来源
PROCEEDINGS OF THE FUTURE TECHNOLOGIES CONFERENCE (FTC) 2021, VOL 2 | 2022年 / 359卷
关键词
Diabetes mellitus; Spearman correlation; Polynomial regression; Random forest; Classification; Machine learning; PIMA Indian; IMPUTATION; TREES;
D O I
10.1007/978-3-030-89880-9_58
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Diabetes mellitus is a metabolic disorder characterized by hyperglycemia which results from the inadequacy of the body to secret and responds to insulin. If not properly managed or diagnosed on time, diabetes can pose a risk to vital body organs such as the eyes, kidneys, nerves, heart, and blood vessels and can be life-threatening. From the many years of research in computational diagnosis of diabetes, machine learning has been proven to be a viable solution for the prediction of diabetes. However, the accuracy rate to date suggests that there is still much room for improvement. In this paper, we are proposing a machine learning framework to improve the performance of diabetes prediction with the PIMA Indian dataset. Through analysis, we observe that the main challenges of the dataset, which flaws learning, are feature selection and missing values. For each of these challenges, we propose a working solution that incorporates, Spearman Correlation and polynomial regression from a new perspective. Further, we optimize the random forest classifier by tuning its hyperparameters using grid search and repeated stratified k-fold cross-validation to build a robust random forest model that scales to the prediction problem. Finally, through exhaustive experiments, we demonstrate that our proposed data preparation approaches lead to a robust machine learning framework for the diagnosis of diabetes mellitus with train accuracy, and test-accuracy values that range from 98.96% to 100% and 97.92% to 100%, respectively, which outperforms all the state-of-the-art results. The source code for the proposed machine learning framework is made publicly available.
引用
收藏
页码:775 / 792
页数:18
相关论文
共 24 条
[1]  
Alam M.T, 2019, INFORM MED UNLOCKED, V16
[2]   Classification and Diagnosis of Diabetes [J].
不详 .
DIABETES CARE, 2015, 38 :S8-S16
[3]  
[Anonymous], 2021, IEEE Trans. Broadcast.
[4]  
Barhate R., 2018, 2018 4 INT C COMPUTI, V4, P1
[5]  
Biau G, 2012, J MACH LEARN RES, V13, P1063
[6]  
Breiman L, 1996, MACH LEARN, V24, P123, DOI 10.1023/A:1018054314350
[7]   Gestational diabetes: diagnosis and management [J].
Cheng, Y. W. ;
Caughey, A. B. .
JOURNAL OF PERINATOLOGY, 2008, 28 (10) :657-664
[8]  
Corder GW., 2014, NONPARAMETRIC STAT S
[9]   Diabetes Prediction Using Ensembling of Different Machine Learning Classifiers [J].
Hasan, Md. Kamrul ;
Alam, Md. Ashraful ;
Das, Dola ;
Hossain, Eklas ;
Hasan, Mahmudul .
IEEE ACCESS, 2020, 8 :76516-76531
[10]   Cross-validation pitfalls when selecting and assessing regression and classification models [J].
Krstajic, Damjan ;
Buturovic, Ljubomir J. ;
Leahy, David E. ;
Thomas, Simon .
JOURNAL OF CHEMINFORMATICS, 2014, 6