Feature selection with missing data using mutual information estimators

被引:58
作者
Doquire, Gauthier [1 ]
Verleysen, Michel [1 ]
机构
[1] Catholic Univ Louvain, Machine Learning Grp, ICTEAM, B-1348 Louvain, Belgium
关键词
Feature selection; Missing data; Mutual information; FUNCTIONAL DATA; VALUES; IMPUTATION; REGRESSION; VARIABLES;
D O I
10.1016/j.neucom.2012.02.031
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Feature selection is an important preprocessing task for many machine learning and pattern recognition applications, including regression and classification. Missing data are encountered in many real-world problems and have to be considered in practice. This paper addresses the problem of feature selection in prediction problems where some occurrences of features are missing. To this end, the well-known mutual information criterion is used. More precisely, it is shown how a recently introduced nearest neighbors based mutual information estimator can be extended to handle missing data. This estimator has the advantage over traditional ones that it does not directly estimate any probability density function. Consequently, the mutual information may be reliably estimated even when the dimension of the space increases. Results on artificial as well as real-world datasets indicate that the method is able to select important features without the need for any imputation algorithm, under the assumption of missing completely at random data. Moreover, experiments show that selecting the features before imputing the data generally increases the precision of the prediction models, in particular when the proportion of missing data is high. (C) 2012 Elsevier B.V. All rights reserved.
引用
收藏
页码:3 / 11
页数:9
相关论文
共 50 条
[31]   Feature Selection with Mutual Information for Regression Problems [J].
Sulaiman, Muhammad Aliyu ;
Labadin, Jane .
2015 9TH INTERNATIONAL CONFERENCE ON IT IN ASIA (CITA), 2015,
[32]   Mutual information for feature selection: estimation or counting? [J].
Nguyen H.B. ;
Xue B. ;
Andreae P. .
Evolutionary Intelligence, 2016, 9 (3) :95-110
[33]   Genetic algorithm for feature selection with mutual information [J].
Ge, Hong ;
Hu, Tianliang .
2014 SEVENTH INTERNATIONAL SYMPOSIUM ON COMPUTATIONAL INTELLIGENCE AND DESIGN (ISCID 2014), VOL 1, 2014, :116-119
[34]   Feature Selection by Maximizing Part Mutual Information [J].
Gao, Wanfu ;
Hu, Liang ;
Zhang, Ping .
2018 INTERNATIONAL CONFERENCE ON SIGNAL PROCESSING AND MACHINE LEARNING (SPML 2018), 2018, :120-127
[35]   A New Approach for Feature Selection from Microarray Data Based on Mutual Information [J].
Tang, Jian ;
Zhou, Shuigeng .
IEEE-ACM TRANSACTIONS ON COMPUTATIONAL BIOLOGY AND BIOINFORMATICS, 2016, 13 (06) :1004-1015
[36]   An Overview of Methods for Feature Selection Based on Mutual Information for Stream Data Classification [J].
Wankhade, Kapil ;
Rane, Dhiraj ;
Thool, Ravindra .
2013 INTERNATIONAL CONFERENCE ON COMMUNICATION SYSTEMS AND NETWORK TECHNOLOGIES (CSNT 2013), 2013, :630-634
[37]   Multi-Objective Feature Selection With Missing Data in Classification [J].
Xue, Yu ;
Tang, Yihang ;
Xu, Xin ;
Liang, Jiayu ;
Neri, Ferrante .
IEEE TRANSACTIONS ON EMERGING TOPICS IN COMPUTATIONAL INTELLIGENCE, 2022, 6 (02) :355-364
[38]   Feature selection for orthogonal broad learning system based on mutual information [J].
Liu, Zhicheng ;
Chen, Bao ;
Xie, Bingxue ;
Qiang, Huangping ;
Zhu, Ziqi .
2019 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS (IJCNN), 2019,
[39]   Mutual information inspired feature selection using kernel canonical correlation analysis [J].
Wang Y. ;
Cang S. ;
Yu H. .
Expert Systems with Applications: X, 2019, 4
[40]   Feature selection using mutual information based uncertainty measures for tumor classification [J].
Sun, Lin ;
Xu, Jiucheng .
BIO-MEDICAL MATERIALS AND ENGINEERING, 2014, 24 (01) :763-770