Improved cost-sensitive representation of data for solving the imbalanced big data classification problem

被引：0

作者：

Mahboubeh Fattahi

Mohammad Hossein Moattar

Yahya Forghani

机构：

[1] Islamic Azad University,Department of Computer Engineering, Mashhad Branch

来源：

Journal of Big Data | / 9卷

关键词：

Feature selection; Feature extraction; Imbalanced data; Big data classification; Cost sensitive; Optimization;

D O I：

暂无

中图分类号：

学科分类号：

摘要：

Dimension reduction is a preprocessing step in machine learning for eliminating undesirable features and increasing learning accuracy. In order to reduce the redundant features, there are data representation methods, each of which has its own advantages. On the other hand, big data with imbalanced classes is one of the most important issues in pattern recognition and machine learning. In this paper, a method is proposed in the form of a cost-sensitive optimization problem which implements the process of selecting and extracting the features simultaneously. The feature extraction phase is based on reducing error and maintaining geometric relationships between data by solving a manifold learning optimization problem. In the feature selection phase, the cost-sensitive optimization problem is adopted based on minimizing the upper limit of the generalization error. Finally, the optimization problem which is constituted from the above two problems is solved by adding a cost-sensitive term to create a balance between classes without manipulating the data. To evaluate the results of the feature reduction, the multi-class linear SVM classifier is used on the reduced data. The proposed method is compared with some other approaches on 21 datasets from the UCI learning repository, microarrays and high-dimensional datasets, as well as imbalanced datasets from the KEEL repository. The results indicate the significant efficiency of the proposed method compared to some similar approaches.

引用

共 49 条

[21]

Shahee SA(2012)SMOTE-RS B*: a hybrid preprocessing approach based on oversampling and undersampling for high imbalanced data-sets using SMOTE and rough sets theory Knowl Inf Syst 224 70-82

[22]

Ananthakumar U(2017)Large cost-sensitive margin distribution machine for imbalanced data classification Neurocomputing 261 214-226

[23]

Chenxi H(2017)Class-specific cost regulation extreme learning machine for imbalanced classification Neurocomputing 200 424-443

[24]

Bennin KE(2020)Joint imbalanced classification and feature selection for hospital readmissions Knowl Based Syst 187 94-105

[25]

Nakariyakul S(2020)SMOTE based class-specific extreme learning machine for imbalanced learning Knowl Based Syst 522 1-23

[26]

Zeng Z(2020)Low-rank matrix regression for image feature extraction and feature selection Inf Sci 25 1-30

[27]

Hart P(2021)Content-based image retrieval based on hybrid feature extraction and feature selection technique pigeon inspired based optimization Ann Roman Soc Cell Biol 67 1-11

[28]

Tomek I(2018)Dealing with high-dimensional class-imbalanced datasets: embedded feature selection for SVM classification Appl Soft Comput 8 7940-7957

[29]

Yen S-J(2021)An alternative approach to dimension reduction for pareto distributed data: a case study J Big Data 7 undefined-undefined

[30]

Lee Y-S(2020)A comprehensive survey of anomaly detection techniques for high dimensional big data J Big Data 4 undefined-undefined

← 1 2 3 4 5 →