A Three-Way Decisions Clustering Algorithm for Incomplete Data

被引:27
作者
Yu, Hong [1 ]
Su, Ting [1 ]
Zeng, Xianhua [1 ]
机构
[1] Chongqing Univ Posts & Telecommun, Chongqing Key Lab Computat Intelligence, Chongqing 400065, Peoples R China
来源
ROUGH SETS AND KNOWLEDGE TECHNOLOGY, RSKT 2014 | 2014年 / 8818卷
关键词
clustering; incomplete data; three-way decisions; attribute significance;
D O I
10.1007/978-3-319-11740-9_70
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Clustering is one of the most widely used efficient approaches in data mining to find potential data structure. However, there are some reasons to cause the missing values in real data sets such as difficulties and limitations of data acquisition and random noises. Most of clustering methods can't be used to deal with incomplete data sets for clustering analysis directly. For this reason, this paper proposes a three-way decisions clustering algorithm for incomplete data based on attribute significance and miss rate. Three-way decisions with interval sets naturally partition a cluster into positive region, boundary region and negative region, which has the advantage of dealing with soft clustering. First, the data set is divided into four parts such as sufficient data, valuable data, inadequate data and invalid data, according to the domain knowledge about the attribute significance and miss rate. Second, different strategies are devised to handle the four types based on three-way decisions. The experimental results on some data sets show preliminarily the effectiveness of the proposed algorithm.
引用
收藏
页码:765 / 776
页数:12
相关论文
共 15 条
[1]  
Azam N, 2013, CAN CON EL COMP EN, P695
[2]  
Dan Li, 2012, Proceedings of the 2012 International Conference on System Science and Engineering (ICSSE), P449, DOI 10.1109/ICSSE.2012.6257226
[3]   PATTERN-RECOGNITION WITH PARTLY MISSING DATA [J].
DIXON, JK .
IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS, 1979, 9 (10) :617-621
[4]  
Himmelspach L, 2011, ADV INTEL SYS RES, P290
[5]  
HONDA K, 2011, 2011 IEEE INT C FUZZ, P1710
[6]  
Hong Yu, 2012, Rough Sets and Current Trends in Computing. Proceedings 8th International Conference, RSCTC 2012, P277, DOI 10.1007/978-3-642-32115-3_33
[7]  
Lai P.-H., 2010, INF THEOR APPL WORKS, P1
[8]   A hybrid genetic algorithm-fuzzy c-means approach for incomplete data clustering based on nearest-neighbor intervals [J].
Li, Dan ;
Gu, Hong ;
Zhang, Liyong .
SOFT COMPUTING, 2013, 17 (10) :1787-1796
[9]  
Liang D.C., 2014, J IEEE T FUZZY SYSTE
[10]   Extended mean field annealing for clustering incomplete data [J].
Wu, Jun ;
Song, Chi-Hwa ;
Kong, Jung Min ;
Lee, Won Don .
2007 INTERNATIONAL SYMPOSIUM ON INFORMATION TECHNOLOGY CONVERGENCE, PROCEEDINGS, 2007, :8-12