Cluster-based under-sampling approaches for imbalanced data distributions

被引:483
|
作者
Yen, Show-Jane
Lee, Yue-Shi
机构
[1] Department of Computer Science and Information Engineering, Ming Chuan University, Gwei Shan District, Taoyuan County 333
关键词
Classification; Data mining; Under-sampling; Imbalanced data distribution;
D O I
10.1016/j.eswa.2008.06.108
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
For classification problem, the training data will significantly influence the classification accuracy. However, data in real-world applications often are imbalanced class distribution, that is, most of the data ever, are in majority class and little data are in minority class. In this case, if all the data are used to be the training data, the classifier tends to predict that most of the incoming data belongs to the majority class. Hence, it is important to select the suitable training data for classification in the imbalanced class distribution problem. In this paper, we propose cluster-based under-sampling approaches for selecting the representative data as training data to improve the classification accuracy for minority class and investigate the effect of under-sampling methods in the imbalanced class distribution environment. The experimental results show that our cluster-based under-sampling approaches outperform the other under-sampling techniques in the previous studies. (C) 2008 Elsevier Ltd. All rights reserved.
引用
收藏
页码:5718 / 5727
页数:10
相关论文
共 50 条
  • [1] Cluster-based sampling approaches to imbalanced data distributions
    Yen, Show-Jane
    Lee, Yue-Shi
    DATA WAREHOUSING AND KNOWLEDGE DISCOVERY, PROCEEDINGS, 2006, 4081 : 427 - 436
  • [2] A Cluster-Based Under-Sampling Algorithm for Class-Imbalanced Data
    Guzman-Ponce, A.
    Valdovinos, R. M.
    Sanchez, J. S.
    HYBRID ARTIFICIAL INTELLIGENT SYSTEMS, HAIS 2020, 2020, 12344 : 299 - 311
  • [3] CUSBoost: Cluster-based Under-sampling with Boosting for Imbalanced Classification
    Rayhan, Farshid
    Ahmed, Sajid
    Mahbub, Asif
    Jani, Md. Rafsan
    Shatabda, Swakkhar
    Farid, Dewan Md.
    2017 2ND INTERNATIONAL CONFERENCE ON COMPUTATIONAL SYSTEMS AND INFORMATION TECHNOLOGY FOR SUSTAINABLE SOLUTION (CSITSS-2017), 2017, : 70 - 75
  • [4] SVM classifier for unbalanced data based on spectrum cluster-based under-sampling approaches
    Tao, Xin-Min
    Zhang, Dong-Xue
    Hao, Si-Yuan
    Fu, Dan-Dan
    Kongzhi yu Juece/Control and Decision, 2012, 27 (12): : 1761 - 1768
  • [5] Cluster-based Majority Under-Sampling Approaches for Class Imbalance Learning
    Zhang, Yan-Ping
    Zhang, Li-Na
    Wang, Yong-Cheng
    2010 2ND IEEE INTERNATIONAL CONFERENCE ON INFORMATION AND FINANCIAL ENGINEERING (ICIFE), 2010, : 400 - 404
  • [6] Cluster-based Under-sampling with Random Forest for Multi-Class Imbalanced Classification
    Arafat, Md. Yasir
    Hoque, Sabera
    Farid, Dewan Md.
    2017 11TH INTERNATIONAL CONFERENCE ON SOFTWARE, KNOWLEDGE, INFORMATION MANAGEMENT AND APPLICATIONS (SKIMA), 2017,
  • [7] Feature Selection and Ensemble Hierarchical Cluster-based Under-sampling Approach for Extremely Imbalanced Datasets
    Soltani, Sima
    Sadri, Javad
    Torshizi, Hassan Ahmadi
    2011 1ST INTERNATIONAL ECONFERENCE ON COMPUTER AND KNOWLEDGE ENGINEERING (ICCKE), 2011, : 166 - 171
  • [8] Cluster-based sampling of multiclass imbalanced data
    Prachuabsupakij, Wanthanee
    Soonthornphisaj, Nuanwan
    INTELLIGENT DATA ANALYSIS, 2014, 18 (06) : 1109 - 1135
  • [9] A Cluster-based Regrouping Approach for Imbalanced Data Distributions
    Yu, Wen
    Jiang, ShengYi
    2012 WORLD AUTOMATION CONGRESS (WAC), 2012,
  • [10] Comparison of Cluster-Based Sampling Approaches for Imbalanced Data of Crashes Involving Large Trucks
    Tahfim, Syed As-Sadeq
    Chen, Yan
    INFORMATION, 2024, 15 (03)