Cluster-based under-sampling approaches for imbalanced data distributions

被引：503

作者：

Yen, Show-Jane

Lee, Yue-Shi

机构：

[1] Department of Computer Science and Information Engineering, Ming Chuan University, Gwei Shan District, Taoyuan County 333

来源：

EXPERT SYSTEMS WITH APPLICATIONS | 2009年 / 36卷 / 03期

关键词：

Classification; Data mining; Under-sampling; Imbalanced data distribution;

D O I：

10.1016/j.eswa.2008.06.108

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

For classification problem, the training data will significantly influence the classification accuracy. However, data in real-world applications often are imbalanced class distribution, that is, most of the data ever, are in majority class and little data are in minority class. In this case, if all the data are used to be the training data, the classifier tends to predict that most of the incoming data belongs to the majority class. Hence, it is important to select the suitable training data for classification in the imbalanced class distribution problem. In this paper, we propose cluster-based under-sampling approaches for selecting the representative data as training data to improve the classification accuracy for minority class and investigate the effect of under-sampling methods in the imbalanced class distribution environment. The experimental results show that our cluster-based under-sampling approaches outperform the other under-sampling techniques in the previous studies. (C) 2008 Elsevier Ltd. All rights reserved.

引用

页码：5718 / 5727

页数：10

共 19 条

[1]

[Anonymous], 2004, ACM SIGKDD EXPLORATI, DOI DOI 10.1145/1007730.1007737

[2]

[Anonymous], 2000, P AAAI 2000 WORKSH L

[3] Committee-based sample selection for probabilistic classifiers [J].

Argamon-Engelson, S ;

Dagan, I .

JOURNAL OF ARTIFICIAL INTELLIGENCE RESEARCH, 1999, 11 :335-360

[4]

Chawla N.V., 2003, P ICML, V3, P66

[5] SMOTE: Synthetic minority over-sampling technique [J].

Chawla, Nitesh V. ;

Bowyer, Kevin W. ;

Hall, Lawrence O. ;

Kegelmeyer, W. Philip .

2002, American Association for Artificial Intelligence (16)

[6] SMOTEBoost: Improving prediction of the minority class in boosting [J].

Chawla, NV ;

Lazarevic, A ;

Hall, LO ;

Bowyer, KW .

KNOWLEDGE DISCOVERY IN DATABASES: PKDD 2003, PROCEEDINGS, 2003, 2838 :107-119

[7]

Chyi Y.M., 2003, THESIS NATL SUN YAT

[8]

del-Hoyo R, 2003, LECT NOTES COMPUT SC, V2686, P334

[9]

Drumnond C., 2003, ICML KDD 2003 WORKSH, V3

[10]

Elkan Charles, 2001, P 17 INT JOINT C ART, V17, P973

← 1 2 →