A variable-selection heuristic for K-means clustering

被引:98
作者
Brusco, MJ [1 ]
Cradit, JD [1 ]
机构
[1] Florida State Univ, Coll Business, Dept Mkt, Tallahassee, FL 32306 USA
关键词
cluster analysis; K-means partitioning; variable selection; heuristics;
D O I
10.1007/BF02294838
中图分类号
O1 [数学];
学科分类号
0701 ; 070101 ;
摘要
One of the most vexing problems in cluster analysis is the selection and/or weighting of variables in order to include those that truly define cluster structure, while eliminating those that might mask such structure. This paper presents a variable-selection heuristic For nonhierarchical (K-means) cluster analysis based on the adjusted Rand index for measuring cluster recovery. The heuristic was subjected to Monte Carlo testing across more than 2200 datasets with known cluster structure. The results indicate the heuristic is extremely effective at eliminating masking variables. A cluster analysis of real-world financial services data revealed that using the variable-selection heuristic prior to the K-means algorithm resulted in greater cluster stability.
引用
收藏
页码:249 / 270
页数:22
相关论文
共 45 条
[1]  
Anderberg M.R., 1973, Probability and Mathematical Statistics
[2]  
Arabie P., 1994, ADV METHODS MARKETIN, P160
[3]  
ART D, 1982, UTILITAS MATHEMATICA, V21, P75
[4]  
BALAKRISHNAN PV, 1994, PSYCHOMETRIKA, V59, P509
[5]   Modeling large data sets in marketing [J].
Balasubramanian, S ;
Gupta, S ;
Kamakura, W ;
Wedel, M .
STATISTICA NEERLANDICA, 1998, 52 (03) :303-323
[6]  
Berry MichaelJ., 1997, DATA MINING TECHNIQU
[7]  
BLATTBERG RC, 1994, MARKETING INFORMATIO
[8]   A NOTE ON THE GENERATION OF RANDOM NORMAL DEVIATES [J].
BOX, GEP ;
MULLER, ME .
ANNALS OF MATHEMATICAL STATISTICS, 1958, 29 (02) :610-611
[9]   REPLICATING CLUSTER-ANALYSIS - METHOD, CONSISTENCY, AND VALIDITY [J].
BRECKENRIDGE, JN .
MULTIVARIATE BEHAVIORAL RESEARCH, 1989, 24 (02) :147-161
[10]   HlNoV: A new model to improve market segment definition by identifying noisy variables [J].
Carmone, FJ ;
Kara, A ;
Maxwell, S .
JOURNAL OF MARKETING RESEARCH, 1999, 36 (04) :501-509