Cluster-based under-sampling approaches for imbalanced data distributions

被引:483
|
作者
Yen, Show-Jane
Lee, Yue-Shi
机构
[1] Department of Computer Science and Information Engineering, Ming Chuan University, Gwei Shan District, Taoyuan County 333
关键词
Classification; Data mining; Under-sampling; Imbalanced data distribution;
D O I
10.1016/j.eswa.2008.06.108
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
For classification problem, the training data will significantly influence the classification accuracy. However, data in real-world applications often are imbalanced class distribution, that is, most of the data ever, are in majority class and little data are in minority class. In this case, if all the data are used to be the training data, the classifier tends to predict that most of the incoming data belongs to the majority class. Hence, it is important to select the suitable training data for classification in the imbalanced class distribution problem. In this paper, we propose cluster-based under-sampling approaches for selecting the representative data as training data to improve the classification accuracy for minority class and investigate the effect of under-sampling methods in the imbalanced class distribution environment. The experimental results show that our cluster-based under-sampling approaches outperform the other under-sampling techniques in the previous studies. (C) 2008 Elsevier Ltd. All rights reserved.
引用
收藏
页码:5718 / 5727
页数:10
相关论文
共 50 条
  • [41] Topographic Under-Sampling for Unbalanced Distributions
    Hamdi, Fatma
    Lebbah, Mustapha
    Bennani, Younes
    2010 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS IJCNN 2010, 2010,
  • [42] A Selective Under-Sampling (SUS) Method For Imbalanced Regression
    Aleksic, Jovana
    Garcia-Remesal, Miguel
    JOURNAL OF ARTIFICIAL INTELLIGENCE RESEARCH, 2025, 82 : 111 - 136
  • [43] Support vector machine for unbalanced data based on sample properties under-sampling approaches
    Tao, X.-M. (taoxinmin@hrbeu.edu.cn), 1600, Northeast University (28):
  • [44] A Selective Under-Sampling based Bagging SVM for Imbalanced Data Learning in Biomedical Event Trigger Recognition
    Chen, Yifei
    2018 2ND INTERNATIONAL CONFERENCE ON BIOMEDICAL ENGINEERING AND BIOINFORMATICS (ICBEB 2018), 2018, : 112 - 119
  • [45] A multi-manifold learning based instance weighting and under-sampling for imbalanced data classification problems
    Tayyebe Feizi
    Mohammad Hossein Moattar
    Hamid Tabatabaee
    Journal of Big Data, 10
  • [46] A multi-manifold learning based instance weighting and under-sampling for imbalanced data classification problems
    Feizi, Tayyebe
    Moattar, Mohammad Hossein
    Tabatabaee, Hamid
    JOURNAL OF BIG DATA, 2023, 10 (01)
  • [47] Framework for the Classification of Imbalanced Structured Data Using Under-sampling and Convolutional Neural Network
    Yoon Sang Lee
    Chulhwan Chris Bang
    Information Systems Frontiers, 2022, 24 : 1795 - 1809
  • [48] A Novel Evolutionary Preprocessing Method Based on Over-sampling and Under-sampling for Imbalanced Datasets
    Wong, Ginny Y.
    Leung, Frank H. F.
    Ling, Sai-Ho
    39TH ANNUAL CONFERENCE OF THE IEEE INDUSTRIAL ELECTRONICS SOCIETY (IECON 2013), 2013, : 2354 - 2359
  • [49] An Under-Sampling Method with Support Vectors in Multi-class Imbalanced Data Classification
    Arafat, Md. Yasir
    Hoque, Sabera
    Xu, Shuxiang
    Farid, Dewan Md.
    2019 13TH INTERNATIONAL CONFERENCE ON SOFTWARE, KNOWLEDGE, INFORMATION MANAGEMENT AND APPLICATIONS (SKIMA), 2019,
  • [50] Two-step ensemble under-sampling algorithm for massive imbalanced data classification
    Bai, Lin
    Ju, Tong
    Wang, Hao
    Lei, Mingzhu
    Pan, Xiaoying
    INFORMATION SCIENCES, 2024, 665