Neural Network-Based Undersampling Techniques

被引:32
|
作者
Arefeen, Md Adnan [1 ,2 ]
Nimi, Sumaiya Tabassum [1 ,2 ]
Rahman, M. Sohel [3 ]
机构
[1] Univ Missouri, Dept Comp Sci Elect Engn, Kansas City, MO 64110 USA
[2] United Int Univ, Dept Comp Sci & Engn, Dhaka 1209, Bangladesh
[3] Bangladesh Univ Engn & Technol, Dept Comp Sci & Engn, Dhaka 1205, Bangladesh
来源
IEEE TRANSACTIONS ON SYSTEMS MAN CYBERNETICS-SYSTEMS | 2022年 / 52卷 / 02期
关键词
Task analysis; Noise measurement; Neurons; Machine learning algorithms; Computer science; Genetic algorithms; Autoencoder; class imbalance; classification; neural network; undersampling; CLASSIFICATION; IMBALANCE; FRAUD; SMOTE;
D O I
10.1109/TSMC.2020.3016283
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Machine learning models have gained popularity nowadays for their potential to solve real-life issues when trained on pertinent data. In many cases, the real-life data are class imbalanced and hence the corresponding machine learning models trained on the data tend to perform poorly on metrics like precision, recall, AUC, F1, and G-mean score. Since class imbalance issue poses serious challenges to the performance of trained models, a multitude of research works have addressed this issue. Two common data-based sampling techniques have mostly been proposed-undersampling the data of the majority class and oversampling the data of the minority class. In this article, we focus on the former approach. We propose two novel algorithms that employ neural network-based approaches to remove majority samples that are found to reside in the vicinity of the minority samples, thereby undersampling the former to remove (or alleviate) the imbalance issue. We delineate the proposed algorithms and then test the proposed algorithms on some publicly available imbalanced datasets. We then compare the performance of our proposed algorithms to other popular undersampling algorithms. Finally, we conclude that our proposed algorithms outperform most of the existing undersampling approaches on most performance metrics.
引用
收藏
页码:1111 / 1120
页数:10
相关论文
共 50 条