Radial-Based oversampling for noisy imbalanced data classification

被引:94
|
作者
Koziarski, Michal [1 ]
Krawczyk, Bartosz [2 ]
Wozniak, Michal [3 ]
机构
[1] AGH Univ Sci & Technol, Dept Elect, Al Mickiewicza 30, PL-30059 Krakow, Poland
[2] Virginia Commonwealth Univ, Dept Comp Sci, 401 West Main St,POB 843019, Richmond, VA 23284 USA
[3] Wroclaw Univ Sci & Technol, Dept Syst & Comp Networks, Wybrzeze Wyspianskiego 27, PL-50370 Wroclaw, Poland
关键词
Pattern classification; Machine learning; Imbalanced data; Oversampling; Radial basis functions; Noisy data; SAMPLING METHOD; MINORITY CLASS; SMOTE; IDENTIFICATION; EXAMPLES;
D O I
10.1016/j.neucom.2018.04.089
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Imbalanced data classification remains a focus of intense research, mostly due to the prevalence of data imbalance in various real-life application domains. A disproportion among objects from different classes may significantly affect the performance of standard classification models. The first problem is the high imbalance ratios that pose a serious learning difficulty and require usage of dedicated methods, capable of alleviating this issue. The second important problem which may appear is noise, which may be accompanying the training data and causing strong deterioration of the classifier performance or increase the time required for its training. Therefore, the desirable classification model should be robust to both skewed data distributions and noise. One of the most popular approaches for handling imbalanced data is oversampling of the minority objects in their neighborhood. In this work we will criticize this approach and propose a novel strategy for dealing with imbalanced data, with particular focus on the noise presence. We propose Radial Based Oversampling (RBO) method, which can find regions in which the synthetic objects from minority class should be generated on the basis of the imbalance distribution estimation with radial basis functions. Results of experiments, carried out on a representative set of benchmark datasets, confirm that the proposed guided synthetic oversampling algorithm offers an interesting alternative to popular state-of-the-art solutions for imbalanced data preprocessing. (C) 2019 Elsevier B.V. All rights reserved.
引用
收藏
页码:19 / 33
页数:15
相关论文
共 50 条
  • [21] Improved KD-tree based imbalanced big data classification and oversampling for MapReduce platforms
    Sleeman, William C.
    Roseberry, Martha
    Ghosh, Preetam
    Cano, Alberto
    Krawczyk, Bartosz
    APPLIED INTELLIGENCE, 2024, 54 (23) : 12558 - 12575
  • [22] A Synthetic Minority Oversampling Technique Based on Gaussian Mixture Model Filtering for Imbalanced Data Classification
    Xu, Zhaozhao
    Shen, Derong
    Kou, Yue
    Nie, Tiezheng
    IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2024, 35 (03) : 3740 - 3753
  • [23] Local distribution-based adaptive minority oversampling for imbalanced data classification
    Wang, Xinyue
    Xu, Jian
    Zeng, Tieyong
    Jing, Liping
    NEUROCOMPUTING, 2021, 422 : 200 - 213
  • [24] A quantum-based oversampling method for classification of highly imbalanced and overlapped data
    Yang, Bei
    Tian, Guilan
    Luttrell, Joseph
    Gong, Ping
    Zhang, Chaoyang
    EXPERIMENTAL BIOLOGY AND MEDICINE, 2023, 248 (24) : 2500 - 2513
  • [25] OVERSAMPLING METHOD FOR IMBALANCED CLASSIFICATION
    Zheng, Zhuoyuan
    Cai, Yunpeng
    Li, Ye
    COMPUTING AND INFORMATICS, 2015, 34 (05) : 1017 - 1037
  • [26] DTO-SMOTE: Delaunay Tessellation Oversampling for Imbalanced Data Sets
    de Carvalho, Alexandre M.
    Prati, Ronaldo C.
    INFORMATION, 2020, 11 (12) : 1 - 22
  • [27] Integrated Oversampling for Imbalanced Time Series Classification
    Cao, Hong
    Li, Xiao-Li
    Woon, David Yew-Kwong
    Ng, See-Kiong
    IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2013, 25 (12) : 2809 - 2822
  • [28] Imbalance: Oversampling algorithms for imbalanced classification in R
    Cordon, Ignacio
    Garcia, Salvador
    Fernandez, Alberto
    Herrera, Francisco
    KNOWLEDGE-BASED SYSTEMS, 2018, 161 : 329 - 341
  • [29] Imbalanced Learning with Oversampling based on Classification Contribution Degree
    Jiang, Zhenhao
    Yang, Jie
    Liu, Yan
    ADVANCED THEORY AND SIMULATIONS, 2021, 4 (05)
  • [30] DDSC-SMOTE: an imbalanced data oversampling algorithm based on data distribution and spectral clustering
    Li, Xinqi
    Liu, Qicheng
    JOURNAL OF SUPERCOMPUTING, 2024, 80 (12) : 17760 - 17789