Cluster-based under-sampling approaches for imbalanced data distributions

被引:483
|
作者
Yen, Show-Jane
Lee, Yue-Shi
机构
[1] Department of Computer Science and Information Engineering, Ming Chuan University, Gwei Shan District, Taoyuan County 333
关键词
Classification; Data mining; Under-sampling; Imbalanced data distribution;
D O I
10.1016/j.eswa.2008.06.108
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
For classification problem, the training data will significantly influence the classification accuracy. However, data in real-world applications often are imbalanced class distribution, that is, most of the data ever, are in majority class and little data are in minority class. In this case, if all the data are used to be the training data, the classifier tends to predict that most of the incoming data belongs to the majority class. Hence, it is important to select the suitable training data for classification in the imbalanced class distribution problem. In this paper, we propose cluster-based under-sampling approaches for selecting the representative data as training data to improve the classification accuracy for minority class and investigate the effect of under-sampling methods in the imbalanced class distribution environment. The experimental results show that our cluster-based under-sampling approaches outperform the other under-sampling techniques in the previous studies. (C) 2008 Elsevier Ltd. All rights reserved.
引用
收藏
页码:5718 / 5727
页数:10
相关论文
共 50 条
  • [21] Automatic incident detection algorithm based on under-sampling for imbalanced traffic data
    Li, Miao-hua
    Chen, Shu-yan
    Lao, Ye-chun
    GREEN BUILDING, ENVIRONMENT, ENERGY AND CIVIL ENGINEERING, 2017, : 145 - 150
  • [22] A Hybrid Under-Sampling Method (HUSBoost) to Classify Imbalanced Data
    Popel, Mahmudul Hasan
    Hasib, Khan Md
    Habib, Syed Ahsan
    Shah, Faisal Muhammad
    2018 21ST INTERNATIONAL CONFERENCE OF COMPUTER AND INFORMATION TECHNOLOGY (ICCIT), 2018,
  • [23] Under-sampling approaches for improving prediction of the minority class in an imbalanced dataset
    Yen, Show-Jane
    Lee, Yue-Shi
    INTELLIGENT CONTROL AND AUTOMATION, 2006, 344 : 731 - 740
  • [24] Using Weight-Retouching and Under-Sampling SVM Approaches for Text Categorization on Imbalanced Data
    Wang He-Yong
    2009 INTERNATIONAL CONFERENCE ON E-BUSINESS AND INFORMATION SYSTEM SECURITY, VOLS 1 AND 2, 2009, : 546 - 549
  • [26] Uncertainty Based Under-Sampling for Learning Naive Bayes Classifiers Under Imbalanced Data Sets
    Aridas, Christos K.
    Karlos, Stamatis
    Kanas, Vasileios G.
    Fazakis, Nikos
    Kotsiantis, Sotiris B.
    IEEE ACCESS, 2020, 8 : 2122 - 2133
  • [27] A design of information granule-based under-sampling method in imbalanced data classification
    Tianyu Liu
    Xiubin Zhu
    Witold Pedrycz
    Zhiwu Li
    Soft Computing, 2020, 24 : 17333 - 17347
  • [28] A design of information granule-based under-sampling method in imbalanced data classification
    Liu, Tianyu
    Zhu, Xiubin
    Pedrycz, Witold
    Li, Zhiwu
    SOFT COMPUTING, 2020, 24 (22) : 17333 - 17347
  • [29] Ensemble based on feature projection and under-sampling for imbalanced learning
    Guo, Huaping
    Zhou, Jun
    Wu, Chang-an
    She, Wei
    Xu, Mingliang
    INTELLIGENT DATA ANALYSIS, 2018, 22 (05) : 959 - 980
  • [30] Multi-granularity relabeled under-sampling algorithm for imbalanced data
    Dai, Qi
    Liu, Jian-wei
    Liu, Yang
    APPLIED SOFT COMPUTING, 2022, 124