Learning Multiple Nonredundant Clusterings

被引：10

作者：

Cui, Ying ^{[1
]}

Fern, Xiaoli Z. ^{[2
]}

Dy, Jennifer G. ^{[3
]}

机构：

[1] Yahoo Inc, Yahoo Labs, Sunnyvale, CA USA

[2] Oregon State Univ, Sch Elect Engn & Comp Sci, Corvallis, OR 97331 USA

[3] Northeastern Univ, Dept Elect & Comp Engn, Boston, MA 02115 USA

来源：

ACM TRANSACTIONS ON KNOWLEDGE DISCOVERY FROM DATA | 2010年 / 4卷 / 03期

关键词：

Nonredundant clustering; disparate clustering; diverse clustering; orthogonalization; FEATURE-SELECTION;

D O I：

10.1145/1839490.1839496

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Real-world applications often involve complex data that can be interpreted in many different ways. When clustering such data, there may exist multiple groupings that are reasonable and interesting from different perspectives. This is especially true for high-dimensional data, where different feature subspaces may reveal different structures of the data. However, traditional clustering is restricted to finding only one single clustering of the data. In this article, we propose a new clustering paradigm for exploratory data analysis: find all non-redundant clustering solutions of the data, where data points in the same cluster in one solution can belong to different clusters in other partitioning solutions. We present a framework to solve this problem and suggest two approaches within this framework: (1) orthogonal clustering, and (2) clustering in orthogonal subspaces. In essence, both approaches find alternative ways to partition the data by projecting it to a space that is orthogonal to the current solution. The first approach seeks orthogonality in the cluster space, while the second approach seeks orthogonality in the feature space. We study the relationship between the two approaches. We also combine our framework with techniques for automatically finding the number of clusters in the different solutions, and study stopping criteria for determining when all meaningful solutions are discovered. We test our framework on both synthetic and high-dimensional benchmark data sets, and the results show that indeed our approaches were able to discover varied clustering solutions that are interesting and meaningful.

引用

页数：32

共 41 条

[1]

Agrawal R., 1998, SIGMOD Record, V27, P94, DOI 10.1145/276305.276314

[2] NEW LOOK AT STATISTICAL-MODEL IDENTIFICATION [J].

AKAIKE, H .

IEEE TRANSACTIONS ON AUTOMATIC CONTROL, 1974, AC19 (06) :716-723

[3]

[Anonymous], 1973, Pattern Classification and Scene Analysis

[4]

Bae E, 2006, IEEE DATA MINING, P53

[5]

BAY S. D, 1999, UCI KDD ACRH

[6]

Blake C. L., 1998, Uci repository of machine learning databases

[7]

Caruana R, 2006, IEEE DATA MINING, P107

[8]

CHECHIK G., 2003, P ADV NEUR INF PROC, P15

[9]

CMU, 1997, CMU 4 U WEBKB DAT

[10]

Cui Y, 2007, IEEE DATA MINING, P133, DOI 10.1109/ICDM.2007.94

← 1 2 3 4 5 →