Radial-Based oversampling for noisy imbalanced data classification

被引:94
|
作者
Koziarski, Michal [1 ]
Krawczyk, Bartosz [2 ]
Wozniak, Michal [3 ]
机构
[1] AGH Univ Sci & Technol, Dept Elect, Al Mickiewicza 30, PL-30059 Krakow, Poland
[2] Virginia Commonwealth Univ, Dept Comp Sci, 401 West Main St,POB 843019, Richmond, VA 23284 USA
[3] Wroclaw Univ Sci & Technol, Dept Syst & Comp Networks, Wybrzeze Wyspianskiego 27, PL-50370 Wroclaw, Poland
关键词
Pattern classification; Machine learning; Imbalanced data; Oversampling; Radial basis functions; Noisy data; SAMPLING METHOD; MINORITY CLASS; SMOTE; IDENTIFICATION; EXAMPLES;
D O I
10.1016/j.neucom.2018.04.089
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Imbalanced data classification remains a focus of intense research, mostly due to the prevalence of data imbalance in various real-life application domains. A disproportion among objects from different classes may significantly affect the performance of standard classification models. The first problem is the high imbalance ratios that pose a serious learning difficulty and require usage of dedicated methods, capable of alleviating this issue. The second important problem which may appear is noise, which may be accompanying the training data and causing strong deterioration of the classifier performance or increase the time required for its training. Therefore, the desirable classification model should be robust to both skewed data distributions and noise. One of the most popular approaches for handling imbalanced data is oversampling of the minority objects in their neighborhood. In this work we will criticize this approach and propose a novel strategy for dealing with imbalanced data, with particular focus on the noise presence. We propose Radial Based Oversampling (RBO) method, which can find regions in which the synthetic objects from minority class should be generated on the basis of the imbalance distribution estimation with radial basis functions. Results of experiments, carried out on a representative set of benchmark datasets, confirm that the proposed guided synthetic oversampling algorithm offers an interesting alternative to popular state-of-the-art solutions for imbalanced data preprocessing. (C) 2019 Elsevier B.V. All rights reserved.
引用
收藏
页码:19 / 33
页数:15
相关论文
共 50 条
  • [1] Radial-Based Oversampling for Multiclass Imbalanced Data Classification
    Krawczyk, Bartosz
    Koziarski, Michal
    Wozniak, Michal
    IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2020, 31 (08) : 2818 - 2831
  • [2] Radial-Based Approach to Imbalanced Data Oversampling
    Koziarski, Michal
    Krawczyk, Bartosz
    Wozniak, Michal
    HYBRID ARTIFICIAL INTELLIGENT SYSTEMS, HAIS 2017, 2017, 10334 : 318 - 327
  • [3] Radial-Based Undersampling for imbalanced data classification
    Koziarski, Michal
    PATTERN RECOGNITION, 2020, 102
  • [4] RB-CCR: Radial-Based Combined Cleaning and Resampling algorithm for imbalanced data classification
    Koziarski, Michal
    Bellinger, Colin
    Wozniak, Michal
    MACHINE LEARNING, 2021, 110 (11-12) : 3059 - 3093
  • [5] RB-CCR: Radial-Based Combined Cleaning and Resampling algorithm for imbalanced data classification
    Michał Koziarski
    Colin Bellinger
    Michał Woźniak
    Machine Learning, 2021, 110 : 3059 - 3093
  • [6] RB-CCR: Radial-Based Combined Cleaning and Resampling algorithm for imbalanced data classification
    Koziarski, Michal
    Bellinger, Colin
    Wozniak, Michal
    2021 IEEE 8TH INTERNATIONAL CONFERENCE ON DATA SCIENCE AND ADVANCED ANALYTICS (DSAA), 2021,
  • [7] Experimental Study on Modified Radial-Based Oversampling
    Bobowska, Barbara
    Wozniak, Michal
    INTERNATIONAL JOINT CONFERENCE SOCO'18-CISIS'18- ICEUTE'18, 2019, 771 : 110 - 119
  • [8] Gaussian Distribution Based Oversampling for Imbalanced Data Classification
    Xie, Yuxi
    Qiu, Min
    Zhang, Haibo
    Peng, Lizhi
    Chen, Zhenxiang
    IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2022, 34 (02) : 667 - 679
  • [9] Radial-based oversampling based on differential evolution for imbalanced dataRadial-based oversampling based on differential...J. Chen et al.
    Jun Chen
    Meng Xia
    Zhijie Wang
    Applied Intelligence, 2025, 55 (7)
  • [10] Adaptive Oversampling for Imbalanced Data Classification
    Ertekin, Seyda
    INFORMATION SCIENCES AND SYSTEMS 2013, 2013, 264 : 261 - 269