Density-based clustering of big probabilistic graphs

被引:18
作者
Halim, Zahid [1 ]
Khattak, Jamal Hussain [1 ,2 ]
机构
[1] Ghulam Ishaq Khan Inst Engn Sci & Technol, Fac Comp Sci & Engn, Topi, Pakistan
[2] Allied Bank Ltd, Informat Technol Grp, Business Solut & Dev, Lahore, Pakistan
关键词
Clustering graphs; Machine learning; Big graphs; Clustering; Community detection; UNCERTAIN DATA; ALGORITHM;
D O I
10.1007/s12530-018-9223-2
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Clustering is a machine learning task to group similar objects in coherent sets. These groups exhibit similar behavior with-in their cluster. With the exponential increase in the data volume, robust approaches are required to process and extract clusters. In addition to large volumes, datasets may have uncertainties due to the heterogeneity of the data sources, resulting in the Big Data. Modern approaches and algorithms in machine learning widely use probability-theory in order to determine the data uncertainty. Such huge uncertain data can be transformed to a probabilistic graph-based representation. This work presents an approach for density-based clustering of big probabilistic graphs. The proposed approach deals with clustering of large probabilistic graphs using the graph's density, where the clustering process is guided by the nodes' degree and the neighborhood information. The proposed approach is evaluated using seven real-world benchmark datasets, namely protein-to-protein interaction, yahoo, movie-lens, core, last.fm, delicious social bookmarking system, and epinions. These datasets are first transformed to a graph-based representation before applying the proposed clustering algorithm. The obtained results are evaluated using three cluster validation indices, namely Davies-Bouldin index, Dunn index, and Silhouette coefficient. This proposal is also compared with four state-of-the-art approaches for clustering large probabilistic graphs. The results obtained using seven datasets and three cluster validity indices suggest better performance of the proposed approach.
引用
收藏
页码:333 / 350
页数:18
相关论文
共 44 条
[1]   A framework for ranking uncertain distributed database [J].
AbdulAzeem, Yousry M. ;
ElDesouky, Ali I. ;
Ali, Hesham A. .
DATA & KNOWLEDGE ENGINEERING, 2014, 92 :1-19
[2]  
Aggarwal CC, 2014, CH CRC DATA MIN KNOW, P1
[3]  
Angelov PY, 2016, IEEE IJCNN, P2405, DOI 10.1109/IJCNN.2016.7727498
[4]  
[Anonymous], ARXIV12105693
[5]  
[Anonymous], CLUST COMPUT
[6]  
[Anonymous], 2011, ADV NEURAL INFORM PR
[7]  
[Anonymous], 2007, Statistics: Cluster Analysis
[8]  
[Anonymous], Em: TKDD, DOI [DOI 10.1109/ICDE.2005.34, 10.1109/ICDE.2005.34]
[9]  
[Anonymous], 2012, P 25 ANN C LEARN THE
[10]   Semantically Enriched Task and Workflow Automation in Crowdsourcing for Linked Data Management [J].
Basharat, Amna ;
Arpinar, I. Budak ;
Dastgheib, Shima ;
Kursuncu, Ugur ;
Kochut, Krys ;
Dogdu, Erdogan .
INTERNATIONAL JOURNAL OF SEMANTIC COMPUTING, 2014, 8 (04) :415-439