Improved cost-sensitive representation of data for solving the imbalanced big data classification problem

被引:0
作者
Mahboubeh Fattahi
Mohammad Hossein Moattar
Yahya Forghani
机构
[1] Islamic Azad University,Department of Computer Engineering, Mashhad Branch
来源
Journal of Big Data | / 9卷
关键词
Feature selection; Feature extraction; Imbalanced data; Big data classification; Cost sensitive; Optimization;
D O I
暂无
中图分类号
学科分类号
摘要
Dimension reduction is a preprocessing step in machine learning for eliminating undesirable features and increasing learning accuracy. In order to reduce the redundant features, there are data representation methods, each of which has its own advantages. On the other hand, big data with imbalanced classes is one of the most important issues in pattern recognition and machine learning. In this paper, a method is proposed in the form of a cost-sensitive optimization problem which implements the process of selecting and extracting the features simultaneously. The feature extraction phase is based on reducing error and maintaining geometric relationships between data by solving a manifold learning optimization problem. In the feature selection phase, the cost-sensitive optimization problem is adopted based on minimizing the upper limit of the generalization error. Finally, the optimization problem which is constituted from the above two problems is solved by adding a cost-sensitive term to create a balance between classes without manipulating the data. To evaluate the results of the feature reduction, the multi-class linear SVM classifier is used on the reduced data. The proposed method is compared with some other approaches on 21 datasets from the UCI learning repository, microarrays and high-dimensional datasets, as well as imbalanced datasets from the KEEL repository. The results indicate the significant efficiency of the proposed method compared to some similar approaches.
引用
收藏
相关论文
共 49 条
[1]  
Rakkeitwinai S(2015)New feature selection for gene expression classification based on degree of class overlap in principal dimensions Comput Biol Med 64 292-298
[2]  
Kabir MM(2011)A new local search based hybrid genetic algorithm for feature selection Neurocomputing 74 2914-2928
[3]  
Shahjahan M(2010)Two cooperative ant colonies for feature selection using fuzzy models Expert Syst Appl 37 2714-2723
[4]  
Murase K(2020)A comprehensive review of dimensionality reduction techniques for feature selection and feature extraction J Appl Sci Technol Trends 1 56-70
[5]  
Vieira SM(2018)A novel efficient feature dimensionality reduction method and its application in engineering Complexity 59 44-58
[6]  
Sousa JM(2020)Overview and comparative study of dimensionality reduction techniques for high dimensional data Inf Fusion 13 571-579
[7]  
Runkler TA(2018)On the role of dimensionality reduction J Comput 5 44-58
[8]  
Zebari R(2017)A novel approach for ontology-based dimensionality reduction for web text document classification Int J Softw Innov 87 154-166
[9]  
Cheng Z(2020)Local feature selection based on artificial immune system for classification Appl Soft Comput 78 717-745
[10]  
Lu Z(2018)Multi-view manifold learning with locality alignment Pattern Recogn 50 280-286