Radial-Based oversampling for noisy imbalanced data classification

被引:94
|
作者
Koziarski, Michal [1 ]
Krawczyk, Bartosz [2 ]
Wozniak, Michal [3 ]
机构
[1] AGH Univ Sci & Technol, Dept Elect, Al Mickiewicza 30, PL-30059 Krakow, Poland
[2] Virginia Commonwealth Univ, Dept Comp Sci, 401 West Main St,POB 843019, Richmond, VA 23284 USA
[3] Wroclaw Univ Sci & Technol, Dept Syst & Comp Networks, Wybrzeze Wyspianskiego 27, PL-50370 Wroclaw, Poland
关键词
Pattern classification; Machine learning; Imbalanced data; Oversampling; Radial basis functions; Noisy data; SAMPLING METHOD; MINORITY CLASS; SMOTE; IDENTIFICATION; EXAMPLES;
D O I
10.1016/j.neucom.2018.04.089
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Imbalanced data classification remains a focus of intense research, mostly due to the prevalence of data imbalance in various real-life application domains. A disproportion among objects from different classes may significantly affect the performance of standard classification models. The first problem is the high imbalance ratios that pose a serious learning difficulty and require usage of dedicated methods, capable of alleviating this issue. The second important problem which may appear is noise, which may be accompanying the training data and causing strong deterioration of the classifier performance or increase the time required for its training. Therefore, the desirable classification model should be robust to both skewed data distributions and noise. One of the most popular approaches for handling imbalanced data is oversampling of the minority objects in their neighborhood. In this work we will criticize this approach and propose a novel strategy for dealing with imbalanced data, with particular focus on the noise presence. We propose Radial Based Oversampling (RBO) method, which can find regions in which the synthetic objects from minority class should be generated on the basis of the imbalance distribution estimation with radial basis functions. Results of experiments, carried out on a representative set of benchmark datasets, confirm that the proposed guided synthetic oversampling algorithm offers an interesting alternative to popular state-of-the-art solutions for imbalanced data preprocessing. (C) 2019 Elsevier B.V. All rights reserved.
引用
收藏
页码:19 / 33
页数:15
相关论文
共 50 条
  • [31] FIAO: Feature Information Aggregation Oversampling for imbalanced data classification
    Wang, Fei
    Zheng, Ming
    Hu, Xiaowen
    Li, Hongchao
    Wang, Taochun
    Chen, Fulong
    APPLIED SOFT COMPUTING, 2024, 161
  • [32] Oversampling Methods for Classification of Imbalanced Breast Cancer Malignancy Data
    Krawczyk, Bartosz
    Jelen, Lukasz
    Krzyzak, Adam
    Fevens, Thomas
    COMPUTER VISION AND GRAPHICS, 2012, 7594 : 483 - 490
  • [33] An Improved D2GAN-based oversampling algorithm for imbalanced data classification
    Zhao, Xiaoqiang
    Yao, Qinglei
    STATISTICAL ANALYSIS AND DATA MINING, 2023, 16 (06) : 569 - 582
  • [34] An intrusion detection imbalanced data classification algorithm based on CWGAN-GP oversampling
    Yao, Qinglei
    Zhao, Xiaoqiang
    PEER-TO-PEER NETWORKING AND APPLICATIONS, 2025, 18 (03)
  • [35] An efficient method to determine sample size in oversampling based on classification complexity for imbalanced data
    Lee, Dohyun
    Kim, Kyoungok
    EXPERT SYSTEMS WITH APPLICATIONS, 2021, 184 (184)
  • [36] A novel oversampling method based on Wasserstein CGAN for imbalanced classification
    Zhou, Hongfang
    Pan, Heng
    Zheng, Kangyun
    Wu, Zongling
    Xiang, Qingyu
    CYBERSECURITY, 2025, 8 (01):
  • [37] A novel oversampling and feature selection hybrid algorithm for imbalanced data classification
    Feng, Fang
    Li, Kuan-Ching
    Yang, Erfu
    Zhou, Qingguo
    Han, Lihong
    Hussain, Amir
    Cai, Mingjiang
    MULTIMEDIA TOOLS AND APPLICATIONS, 2023, 82 (03) : 3231 - 3267
  • [38] Classification of Imbalanced Data by Oversampling in Kernel Space of Support Vector Machines
    Mathew, Josey
    Pang, Chee Khiang
    Luo, Ming
    Leong, Weng Hoe
    IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2018, 29 (09) : 4065 - 4076
  • [39] Evidence-based adaptive oversampling algorithm for imbalanced classification
    Lin, Chen-ju
    Leony, Florence
    KNOWLEDGE AND INFORMATION SYSTEMS, 2024, 66 (03) : 2209 - 2233
  • [40] Perturbation-based oversampling technique for imbalanced classification problems
    Zhang, Jianjun
    Wang, Ting
    Ng, Wing W. Y.
    Pedrycz, Witold
    INTERNATIONAL JOURNAL OF MACHINE LEARNING AND CYBERNETICS, 2023, 14 (03) : 773 - 787