An efficient method to determine sample size in oversampling based on classification complexity for imbalanced data

被引:17
|
作者
Lee, Dohyun [1 ]
Kim, Kyoungok [2 ]
机构
[1] Seoul Natl Univ Sci & Technol Seoul, Dept Data Sci, 232 Gongreungno, Seoul 01811, South Korea
[2] Seoul Natl Univ Sci & Technol Seoul, Dept Ind Engn, 232 Gongreungno, Seoul 01811, South Korea
基金
新加坡国家研究基金会;
关键词
Class imbalance; Oversampling; Sampling size; Adaptive boosting; Ensemble learning; DATA-SETS; SMOTE; ENSEMBLES;
D O I
10.1016/j.eswa.2021.115442
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Resampling, one of the approaches to handle class imbalance, is widely used alone or in combination with other approaches, such as cost-sensitive learning and ensemble learning because of its simplicity and independence in learning algorithms. Oversampling methods, in particular, alleviate class imbalance by increasing the size of the minority class. However, previous studies related to oversampling generally have focused on where to add new samples, how to generate new samples, and how to prevent noise and they rarely have investigated how much sampling is sufficient. In many cases, the oversampling size is set so that the minority class has the same size as the majority class. This setting only considers the size of the classes in sample size determination, and the balanced training set can induce overfitting with the addition of too many minority samples. Moreover, the effectiveness of oversampling can be improved by adding synthetics into the appropriate locations. To address this issue, this study proposes a method to determine the oversampling size less than the sample size needed to obtain a balance between classes, while considering not only the absolute imbalance but also the difficulty of classification in a dataset on the basis of classification complexity. The effectiveness of the proposed sample size in oversampling is evaluated using several boosting algorithms with different oversampling methods for 16 imbalanced datasets. The results show that the proposed sample size achieves better classification performance than the sample size for attaining class balance.
引用
收藏
页数:10
相关论文
共 50 条
  • [41] Stochastic Sensitivity Measure-based Noise Filtering and Oversampling Method for Imbalanced Classification Problems
    Zhang, Jianjun
    Ng, Wing W. Y.
    2018 IEEE INTERNATIONAL CONFERENCE ON SYSTEMS, MAN, AND CYBERNETICS (SMC), 2018, : 403 - 408
  • [42] A new instance density-based synthetic minority oversampling method for imbalanced classification problems
    Ma, Chung-Kang
    Park, You-Jin
    ENGINEERING OPTIMIZATION, 2022, 54 (10) : 1743 - 1757
  • [43] MKC-SMOTE: A Novel Synthetic Oversampling Method for Multi-Class Imbalanced Data Classification
    Wang, Jiao
    Awang, Norhashidah
    IEEE ACCESS, 2024, 12 : 196929 - 196938
  • [44] Adaptive Fusion Based Method for Imbalanced Data Classification
    Liang, Zefeng
    Wang, Huan
    Yang, Kaixiang
    Shi, Yifan
    FRONTIERS IN NEUROROBOTICS, 2022, 16
  • [45] A new boundary-degree-based oversampling method for imbalanced data
    Yueqi Chen
    Witold Pedrycz
    Jie Yang
    Applied Intelligence, 2023, 53 : 26518 - 26541
  • [46] Imbalanced data classification based on diverse sample generation and classifier fusion
    Zhai, Junhai
    Qi, Jiaxing
    Zhang, Sufang
    INTERNATIONAL JOURNAL OF MACHINE LEARNING AND CYBERNETICS, 2022, 13 (03) : 735 - 750
  • [47] Model-Based Oversampling for Imbalanced Sequence Classification
    Gong, Zhichen
    Chen, Huanhuan
    CIKM'16: PROCEEDINGS OF THE 2016 ACM CONFERENCE ON INFORMATION AND KNOWLEDGE MANAGEMENT, 2016, : 1009 - 1018
  • [48] SMOTE-BD: An Exact and Scalable Oversampling Method for Imbalanced Classification in Big Data
    Basgall, Maria Jose
    Hasperue, Waldo
    Naiouf, Marcelo
    Fernandez, Alberto
    Herrera, Francisco
    JOURNAL OF COMPUTER SCIENCE & TECHNOLOGY, 2018, 18 (03): : 203 - 209
  • [49] An oversampling method for wafer map defect pattern classification considering small and imbalanced data
    Kim, Eun-Su
    Choi, Seung-Hyun
    Lee, Dong-Hee
    Kim, Kwang-Jae
    Bae, Young-Mok
    Oh, Young-Chan
    COMPUTERS & INDUSTRIAL ENGINEERING, 2021, 162
  • [50] Distance-based arranging oversampling technique for imbalanced data
    Dai, Qi
    Liu, Jian-wei
    Zhao, Jia-Liang
    NEURAL COMPUTING & APPLICATIONS, 2023, 35 (02) : 1323 - 1342