Prediction of diabetes disease using an ensemble of machine learning multi-classifier models

被引:34
作者
Abnoosian, Karlo [1 ]
Farnoosh, Rahman [2 ]
Behzadi, Mohammad Hassan [1 ]
机构
[1] Islamic Azad Univ, Dept Stat, Sci & Res Branch, Tehran, Iran
[2] Iran Univ Sci & Technol, Sch Math, Tehran, Iran
关键词
Diabetes disease prediction; Machine learning classifiers; Ensemble machine learning models; Decision tree; Random forest; Feature selection; FEATURE-SELECTION;
D O I
10.1186/s12859-023-05465-z
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Background and objectiveDiabetes is a life-threatening chronic disease with a growing global prevalence, necessitating early diagnosis and treatment to prevent severe complications. Machine learning has emerged as a promising approach for diabetes diagnosis, but challenges such as limited labeled data, frequent missing values, and dataset imbalance hinder the development of accurate prediction models. Therefore, a novel framework is required to address these challenges and improve performance.MethodsIn this study, we propose an innovative pipeline-based multi-classification framework to predict diabetes in three classes: diabetic, non-diabetic, and prediabetes, using the imbalanced Iraqi Patient Dataset of Diabetes. Our framework incorporates various pre-processing techniques, including duplicate sample removal, attribute conversion, missing value imputation, data normalization and standardization, feature selection, and k-fold cross-validation. Furthermore, we implement multiple machine learning models, such as k-NN, SVM, DT, RF, AdaBoost, and GNB, and introduce a weighted ensemble approach based on the Area Under the Receiver Operating Characteristic Curve (AUC) to address dataset imbalance. Performance optimization is achieved through grid search and Bayesian optimization for hyper-parameter tuning.ResultsOur proposed model outperforms other machine learning models, including k-NN, SVM, DT, RF, AdaBoost, and GNB, in predicting diabetes. The model achieves high average accuracy, precision, recall, F1-score, and AUC values of 0.9887, 0.9861, 0.9792, 0.9851, and 0.999, respectively.ConclusionOur pipeline-based multi-classification framework demonstrates promising results in accurately predicting diabetes using an imbalanced dataset of Iraqi diabetic patients. The proposed framework addresses the challenges associated with limited labeled data, missing values, and dataset imbalance, leading to improved prediction performance. This study highlights the potential of machine learning techniques in diabetes diagnosis and management, and the proposed framework can serve as a valuable tool for accurate prediction and improved patient care. Further research can build upon our work to refine and optimize the framework and explore its applicability in diverse datasets and populations.
引用
收藏
页数:24
相关论文
共 67 条
[1]  
Abbas NAM., 2020, IRAQI J ELECT ELECT, V16, P1
[2]   Machine learning-based heart disease diagnosis: A systematic literature review [J].
Ahsan, Md Manjurul ;
Siddique, Zahed .
ARTIFICIAL INTELLIGENCE IN MEDICINE, 2022, 128
[3]  
Ali P.J.M., 2014, Mach Learn Tech Rep, V1, P1, DOI [10.13140/RG.2.2.28948.04489, DOI 10.13140/RG.2.2.28948.04489]
[4]  
Anguita D., 2012, InESANN, P441
[5]   Multi-classification by using tri-class SVM [J].
Angulo, C ;
Ruiz, FJ ;
González, L ;
Ortega, JA .
NEURAL PROCESSING LETTERS, 2006, 23 (01) :89-101
[6]   Random forest in remote sensing: A review of applications and future directions [J].
Belgiu, Mariana ;
Dragut, Lucian .
ISPRS JOURNAL OF PHOTOGRAMMETRY AND REMOTE SENSING, 2016, 114 :24-31
[7]   Diagnosed Chronic Health Conditions Among Injured Workers With Permanent Impairments and the General Population [J].
Casey, Rebecca ;
Ballantyne, Peri J. .
JOURNAL OF OCCUPATIONAL AND ENVIRONMENTAL MEDICINE, 2017, 59 (05) :486-496
[8]  
Changsheng Zhu, 2019, Informatics in Medicine Unlocked, V17, P19, DOI 10.1016/j.imu.2019.100179
[9]  
Charbuty B., 2021, Journal of Applied Science and Technology Trends, DOI DOI 10.38094/JASTT20165
[10]   Clinical Review of Antidiabetic Drugs: Implications for Type 2 Diabetes Mellitus Management [J].
Chaudhury, Arun ;
Duvoor, Chitharanjan ;
Dendi, Vijaya Sena Reddy ;
Kraleti, Shashank ;
Chada, Aditya ;
Ravilla, Rahul ;
Marco, Asween ;
Shekhawat, Nawal Singh ;
Montales, Maria Theresa ;
Kuriakose, Kevin ;
Sasapu, Appalanaidu ;
Beebe, Alexandria ;
Patil, Naveen ;
Musham, Chaitanya K. ;
Lohani, Govinda Prasad ;
Mirza, Wasique .
FRONTIERS IN ENDOCRINOLOGY, 2017, 8