Missing data imputation with fuzzy feature selection for diabetes dataset

被引:0
作者
Mohamad Faiz Dzulkalnine
Roselina Sallehuddin
机构
[1] Universiti Teknologi Malaysia,Faculty of Computing
来源
SN Applied Sciences | 2019年 / 1卷
关键词
Missing data; Fuzzy feature selection; Imputation; Classification;
D O I
暂无
中图分类号
学科分类号
摘要
Missing data in datasets remain as a difficulty in terms of data analysis in various research fields, especially in the medical field, as it affects the treatment and diagnosis that the patient should receive. In this research, Fuzzy c-means (FCM) are used to impute the missing data. However, like in most data imputation methods, FCM do not consider the presence of irrelevant features. Irrelevant features can increase the computational time of the imputation process and decrease the accuracy of the prediction. Feature selection techniques can alleviate this problem by selecting the most relevant features and reducing the dataset size. Fuzzy principal component analysis (FPCA) is used as the feature selection method in this study as it considers the presence of outliers compared to classical PCA as outliers are the main reason some features renders irrelevant. Therefore, an improved hybrid imputation model of FPCA–Support vector machines–FCM (FPCA–SVM–FCM) has been proposed and employed in this study. The efficiency of the proposed model is investigated on one dataset which is Pima Indians Diabetes dataset. Experimental results showed that the proposed hybrid imputation model is better than the existing methods by producing a more accurate estimation in terms of accuracy, RMSE and MAE. The proposed method was also validated by using Wilcoxon rank sum and Theil’s U test and obtained good results compared to SVM–FCM. Therefore, it can be used as an alternative tool for handling missing data in order to obtain a better quality dataset.
引用
收藏
相关论文
共 50 条
[31]   A systematic review of generative adversarial imputation network in missing data imputation [J].
Yuqing Zhang ;
Runtong Zhang ;
Butian Zhao .
Neural Computing and Applications, 2023, 35 :19685-19705
[32]   Fuzzy C-mean Missing Data Imputation for Analogy-based Effort Estimation [J].
AlMutlaq, Ayman Jalal ;
Jawawi, Dayang N. A. ;
Arbain, Adila Firdaus Binti .
INTERNATIONAL JOURNAL OF ADVANCED COMPUTER SCIENCE AND APPLICATIONS, 2021, 12 (08) :628-640
[33]   A systematic review of generative adversarial imputation network in missing data imputation [J].
Zhang, Yuqing ;
Zhang, Runtong ;
Zhao, Butian .
NEURAL COMPUTING & APPLICATIONS, 2023, 35 (27) :19685-19705
[34]   Four Factors Affecting Missing Data Imputation [J].
Hackl, Andreas ;
Zeindl, Juergen ;
Ehrlinger, Lisa .
35TH INTERNATIONAL CONFERENCE ON SCIENTIFIC AND STATISTICAL DATABASE MANAGEMENT, SSDBM 2023, 2023,
[35]   Imputation of missing longitudinal data: a comparison of methods [J].
Engels, JM ;
Diehr, P .
JOURNAL OF CLINICAL EPIDEMIOLOGY, 2003, 56 (10) :968-976
[36]   Imputation of Missing Diagnosis of Diabetes in an Administrative EMR System [J].
Wang, Debby D. ;
Ng, Sheryl Hui-Xian ;
Abdul, Siti Nabilah Binte ;
Ramachandran, Sravan ;
Sridharan, Srinath ;
Tan, Xin Quan .
2018 11TH BIOMEDICAL ENGINEERING INTERNATIONAL CONFERENCE (BMEICON 2018), 2018,
[37]   Imputation of missing information in worldwide patent data [J].
de Rassenfosse, Gaetan ;
Seliger, Florian .
DATA IN BRIEF, 2021, 34
[38]   Some Imputation Algorithms for Restoration of Missing Data [J].
Ryazanov, Vladimir .
PROGRESS IN PATTERN RECOGNITION, IMAGE ANALYSIS, COMPUTER VISION, AND APPLICATIONS, 2011, 7042 :372-379
[39]   Missing Data Imputation: A Survey [J].
Kelkar, Bhagyashri Abhay .
INTERNATIONAL JOURNAL OF DECISION SUPPORT SYSTEM TECHNOLOGY, 2022, 14 (01)
[40]   A MISSING DATA IMPUTATION METHOD WITH DISTANCE FUNCTION [J].
Jea, Kuen-Fang ;
Hsu, Chin-Wei ;
Tang, Li-You .
PROCEEDINGS OF 2018 INTERNATIONAL CONFERENCE ON MACHINE LEARNING AND CYBERNETICS (ICMLC), VOL 2, 2018, :450-455