An Empirical Evaluation of Feature Selection Stability and Classification Accuracy

被引:0
|
作者
Buyukkececi, Mustafa [1 ]
Okur, Mehmet Cudi [2 ]
机构
[1] Univerlist, Izmir, Turkiye
[2] Yasar Univ, Fac Engn, Dept Software Engn, Izmir, Turkiye
来源
GAZI UNIVERSITY JOURNAL OF SCIENCE | 2024年 / 37卷 / 02期
关键词
Feature selection; Selection stability; Classification accuracy; Filter methods; Wrapper methods; ALGORITHMS; BIAS;
D O I
10.35378/gujs.998964
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
The performance of inductive learners can be negatively affected by high -dimensional datasets. To address this issue, feature selection methods are used. Selecting relevant features and reducing data dimensions is essential for having accurate machine learning models. Stability is an important criterion in feature selection. Stable feature selection algorithms maintain their feature preferences even when small variations exist in the training set. Studies have emphasized the importance of stable feature selection, particularly in cases where the number of samples is small and the dimensionality is high. In this study, we evaluated the relationship between stability measures, as well as, feature selection stability and classification accuracy, using the Pearson 's Correlation Coefficient (also known as Pearson 's Product -Moment Correlation Coefficient or simply Pearson's r ). We conducted an extensive series of experiments using five filter and two wrapper feature selection methods, three classifiers for subset and classification performance evaluation, and eight real -world datasets taken from two different data repositories. We measured the stability of feature selection methods using a total of twelve stability metrics. Based on the results of correlation analyses, we have found that there is a lack of substantial evidence supporting a linear relationship between feature selection stability and classification accuracy. However, a strong positive correlation has been observed among several stability metrics.
引用
收藏
页码:606 / 620
页数:15
相关论文
共 50 条
  • [21] On the Stability of Feature Selection in the Presence of Feature Correlations
    Sechidis, Konstantinos
    Papangelou, Konstantinos
    Nogueira, Sarah
    Weatherall, James
    Brown, Gavin
    MACHINE LEARNING AND KNOWLEDGE DISCOVERY IN DATABASES, ECML PKDD 2019, PT I, 2020, 11906 : 327 - 342
  • [22] The Effect of Feature Selection on Phish Website Detection An Empirical Study on Robust Feature Subset Selection for Effective Classification
    Zuhair, Hiba
    Selmat, Ali
    Salleh, Mazleena
    INTERNATIONAL JOURNAL OF ADVANCED COMPUTER SCIENCE AND APPLICATIONS, 2015, 6 (10) : 221 - 232
  • [23] Feature Evaluation and Selection for Polarimetric SAR Image Classification
    Chen, Lijun
    Yang, Wen
    Liu, Ying
    Sun, Hong
    2010 IEEE 10TH INTERNATIONAL CONFERENCE ON SIGNAL PROCESSING PROCEEDINGS (ICSP2010), VOLS I-III, 2010, : 2202 - 2205
  • [24] Feature selection algorithms in classification problems: an experimental evaluation
    Salappa, A.
    Doumpos, M.
    Zopounidis, C.
    OPTIMIZATION METHODS & SOFTWARE, 2007, 22 (01) : 199 - 214
  • [25] Using Dimension Reduction with Feature Selection to Enhance Accuracy of Tumor Classification
    Thuy Hang Dang
    Trung Dung Pham
    Hoai Linh Tran
    Quang Lc Van
    2016 3RD INTERNATIONAL CONFERENCE ON BIOMEDICAL ENGINEERING (BME-HUST), 2016, : 14 - 17
  • [26] Accuracy Enhancement for Breast Cancer Detection Using Classification and Feature Selection
    Jain, Somil
    Kumar, Puneet
    INTERNATIONAL JOURNAL OF INFORMATION RETRIEVAL RESEARCH, 2022, 12 (02)
  • [27] An Empirical Study on the Performance of Rule-Based Classification by Feature Selection
    Balakrishnan, Sarojini
    Babu, M. R.
    Krishna, P. V.
    2014 WORLD CONGRESS ON COMPUTING AND COMMUNICATION TECHNOLOGIES (WCCCT 2014), 2014, : 147 - +
  • [28] Empirical Evaluation of the Performance of Feature Selection Approaches on Random Forest
    Kumar, Smitha S.
    Shaikh, Talal
    2017 INTERNATIONAL CONFERENCE ON COMPUTER AND APPLICATIONS (ICCA), 2017, : 227 - 231
  • [29] Empirical Evaluation of the Ensemble Framework for Feature Selection in DDoS Attack
    Das, Saikat
    Venugopal, Deepak
    Shiva, Sajjan
    Sheldon, Frederick T.
    2020 7TH IEEE INTERNATIONAL CONFERENCE ON CYBER SECURITY AND CLOUD COMPUTING (CSCLOUD 2020)/2020 6TH IEEE INTERNATIONAL CONFERENCE ON EDGE COMPUTING AND SCALABLE CLOUD (EDGECOM 2020), 2020, : 56 - 61
  • [30] An Empirical Study on the Stability of Feature Selection for Imbalanced Software Engineering Data
    Wang, Huanjing
    Khoshgoftaar, Taghi M.
    Napolitano, Amri
    2012 11TH INTERNATIONAL CONFERENCE ON MACHINE LEARNING AND APPLICATIONS (ICMLA 2012), VOL 1, 2012, : 317 - 323