Effects of Resampling in Determining the Number of Clusters in a Data Set

被引：0

作者：

Rainer Dangl

Friedrich Leisch

机构：

[1] University of Natural Resources and Life Sciences,Institute for Applied Statistics and Computing

来源：

Journal of Classification | 2020年 / 37卷

关键词：

Resampling; Model validation; Cluster stability; Clustering; Benchmarking;

D O I：

暂无

中图分类号：

学科分类号：

摘要：

Using cluster validation indices is a widely applied method in order to detect the number of groups in a data set and as such a crucial step in the model validation process in clustering. The study presented in this paper demonstrates how the accuracy of certain indices can be significantly improved when calculated numerous times on data sets resampled from the original data. There are obviously many ways to resample data—in this study, three very common options are used: bootstrapping, data splitting (without subset overlap of two subsamples), and random subsetting (with subset overlap of two subsamples). Index values calculated on the basis of resampled data sets are compared to the values obtained from the original data partition. The primary hypothesis of the study states that resampling does generally improve index accuracy. The hypothesis is based on the notion of cluster stability: if there are stable clusters in a data set, a clustering algorithm should produce consistent results for data sampled or resampled from the same source. The primary hypothesis was partly confirmed; for external validation measures, it does indeed apply. The secondary hypothesis states that the resampling strategy itself does not play a significant role. This was also shown to be accurate, yet slight deviations between the resampling schemes suggest that splitting appears to yield slightly better results.

引用

页码：558 / 583

页数：25

共 50 条

[1] Effects of Resampling in Determining the Number of Clusters in a Data Set
Dangl, Rainer
Leisch, Friedrich
JOURNAL OF CLASSIFICATION, 2020, 37 (03) : 558 - 583
[2] Nbclust: An R Package for Determining the Relevant Number of Clusters in a Data Set
Charrad, Malika
Ghazzali, Nadia
Boiteau, Veronique
Niknafs, Azam
JOURNAL OF STATISTICAL SOFTWARE, 2014, 61 (06): : 1 - 36
[3] Automatically Determining the Number of Clusters in Unlabeled Data Sets
Wang, Liang
Leckie, Christopher
Ramamohanarao, Kotagiri
Bezdek, James
IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2009, 21 (03) : 335 - 350
[4] Determining the number of clusters using information entropy for mixed data
Liang, Jiye
Zhao, Xingwang
Li, Deyu
Cao, Fuyuan
Dang, Chuangyin
PATTERN RECOGNITION, 2012, 45 (06) : 2251 - 2265
[5] Automatically Determining the Number of Clusters Using Decision-Theoretic Rough Set
Yu, Hong
Liu, Zhanguo
Wang, Guoyin
ROUGH SETS AND KNOWLEDGE TECHNOLOGY, 2011, 6954 : 504 - 513
[6] Estimating the number of clusters in a data set via the gap statistic
Tibshirani, R
Walther, G
Hastie, T
JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY, 2001, 63 : 411 - 423
[7] Local and Global Data Spread Based Index for Determining Number of Clusters in a Dataset
Riyaz, Romana
Wani, M. Arif
2016 15TH IEEE INTERNATIONAL CONFERENCE ON MACHINE LEARNING AND APPLICATIONS (ICMLA 2016), 2016, : 651 - 656
[8] Estimating the number of clusters in a numerical data set via quantization error modeling
Kolesnikov, Alexander
Trichina, Elena
Kauranne, Tuomo
PATTERN RECOGNITION, 2015, 48 (03) : 941 - 952
[9] DETERMINING THE OPTIMAL NUMBER OF CLUSTERS IN CLUSTER ANALYSIS
Loster, Tomas
10TH INTERNATIONAL DAYS OF STATISTICS AND ECONOMICS, 2016, : 1078 - 1090
[10] EVALUATION OF COEFFICIENTS FOR DETERMINING THE OPTIMAL NUMBER OF CLUSTERS IN CLUSTER ANALYSIS ON REAL DATA SETS
Loster, Tomas
9TH INTERNATIONAL DAYS OF STATISTICS AND ECONOMICS, 2015, : 1014 - 1023

← 1 2 3 4 5 →