An Empirical Evaluation of Feature Selection Stability and Classification Accuracy

被引:0
|
作者
Buyukkececi, Mustafa [1 ]
Okur, Mehmet Cudi [2 ]
机构
[1] Univerlist, Izmir, Turkiye
[2] Yasar Univ, Fac Engn, Dept Software Engn, Izmir, Turkiye
来源
GAZI UNIVERSITY JOURNAL OF SCIENCE | 2024年 / 37卷 / 02期
关键词
Feature selection; Selection stability; Classification accuracy; Filter methods; Wrapper methods; ALGORITHMS; BIAS;
D O I
10.35378/gujs.998964
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
The performance of inductive learners can be negatively affected by high -dimensional datasets. To address this issue, feature selection methods are used. Selecting relevant features and reducing data dimensions is essential for having accurate machine learning models. Stability is an important criterion in feature selection. Stable feature selection algorithms maintain their feature preferences even when small variations exist in the training set. Studies have emphasized the importance of stable feature selection, particularly in cases where the number of samples is small and the dimensionality is high. In this study, we evaluated the relationship between stability measures, as well as, feature selection stability and classification accuracy, using the Pearson 's Correlation Coefficient (also known as Pearson 's Product -Moment Correlation Coefficient or simply Pearson's r ). We conducted an extensive series of experiments using five filter and two wrapper feature selection methods, three classifiers for subset and classification performance evaluation, and eight real -world datasets taken from two different data repositories. We measured the stability of feature selection methods using a total of twelve stability metrics. Based on the results of correlation analyses, we have found that there is a lack of substantial evidence supporting a linear relationship between feature selection stability and classification accuracy. However, a strong positive correlation has been observed among several stability metrics.
引用
收藏
页码:606 / 620
页数:15
相关论文
共 50 条
  • [41] Performance Evaluation of Feature Selection Algorithms on Human Activity Classification
    Tulum, Gokalp
    Artug, N. Tugrul
    Bolat, Bulent
    2013 IEEE INTERNATIONAL SYMPOSIUM ON INNOVATIONS IN INTELLIGENT SYSTEMS AND APPLICATIONS (IEEE INISTA), 2013,
  • [42] An Evaluation on the Efficiency of Hybrid Feature Selection in Spam Email Classification
    Mohamad, Masurah
    Selamat, Ali
    2015 2ND INTERNATIONAL CONFERENCE ON COMPUTER, COMMUNICATIONS, AND CONTROL TECHNOLOGY (I4CT), 2015,
  • [43] A COMPREHENSIVE EVALUATION OF FEATURE SELECTION ALGORITHMS IN HYPERSPECTRAL IMAGE CLASSIFICATION
    Vijouyeh, Hamed G.
    Taskin, Gulsen
    2016 IEEE INTERNATIONAL GEOSCIENCE AND REMOTE SENSING SYMPOSIUM (IGARSS), 2016, : 489 - 492
  • [44] An empirical study on the joint impact of feature selection and data resampling on imbalance classification
    Chongsheng Zhang
    Paolo Soda
    Jingjun Bi
    Gaojuan Fan
    George Almpanidis
    Salvador García
    Weiping Ding
    Applied Intelligence, 2023, 53 : 5449 - 5461
  • [45] Local Causal and Markov Blanket Induction for Causal Discovery and Feature Selection for Classification Part I: Algorithms and Empirical Evaluation
    Aliferis, Constantin F.
    Statnikov, Alexander
    Tsamardinos, Ioannis
    Mani, Subramani
    Koutsoukos, Xenofon D.
    JOURNAL OF MACHINE LEARNING RESEARCH, 2010, 11 : 171 - 234
  • [46] Evaluation of feature selection techniques on network traffic for comparing model accuracy
    Kaur, Prabhjot
    Awasthi, Amit
    Bijalwan, Anchit
    INTERNATIONAL JOURNAL OF COMPUTATIONAL SCIENCE AND ENGINEERING, 2021, 24 (03) : 228 - 243
  • [47] Stability of Feature Selection Algorithms and its Influence on Prediction Accuracy in Biomedical Datasets
    Drotar, Peter
    Smekal, Zdenek
    TENCON 2014 - 2014 IEEE REGION 10 CONFERENCE, 2014,
  • [48] Feature selection using dynamic weights for classification
    Sun, Xin
    Liu, Yanheng
    Xu, Mantao
    Chen, Huiling
    Han, Jiawei
    Wang, Kunhao
    KNOWLEDGE-BASED SYSTEMS, 2013, 37 : 541 - 549
  • [49] A unified pipeline for online feature selection and classification
    Bolon-Canedo, Veronica
    Fernandez-Francos, Diego
    Peteiro-Barral, Diego
    Alonso-Betanzos, Amparo
    Guijarro-Berdinas, Bertha
    Sanchez-Marono, Noelia
    EXPERT SYSTEMS WITH APPLICATIONS, 2016, 55 : 532 - 545
  • [50] Feature Selection for Classification of Hyperspectral Data by SVM
    Pal, Mahesh
    Foody, Giles M.
    IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 2010, 48 (05): : 2297 - 2307