A Comprehensive Survey on Cloud Data Mining (CDM) Frameworks and Algorithms

被引:17
|
作者
Barua, Hrishav Bakul [1 ,3 ]
Mondal, Kartick Chandra [2 ]
机构
[1] Embedded Syst & Robot Res Grp, TCS Res & Innovat Lab, Kolkata, India
[2] Jadavpur Univ, Dept Informat Technol, Sect 3, Kolkata 700106, W Bengal, India
[3] TCS Ecospace, TCS Res & Innovat Lab, Act Area 2, Kolkata 700156, W Bengal, India
关键词
Review; survey; taxonomy; framework; data mining; machine learning; distributed computing; cloud data mining (CDM); big data; big data analytics; data science; cloud computing; parallelism; graph mining; volume; velocity; variety; clustering; classification and association rule mining; BIG DATA; CLUSTERING ALGORITHMS; DATA ANALYTICS; MAPREDUCE; DBSCAN; CLASSIFICATION; PRIVACY; STORAGE; RISE;
D O I
10.1145/3349265
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Data mining is used for finding meaningful information out of a vast expanse of data. With the advent of Big Data concept, data mining has come to much more prominence. Discovering knowledge out of a gigantic volume of data efficiently is a major concern as the resources are limited. Cloud computing plays a major role in such a situation. Cloud data mining fuses the applicability of classical data mining with the promises of cloud computing. This allows it to perform knowledge discovery out of huge volumes of data with efficiency. This article presents the existing frameworks, services, platforms, and algorithms for cloud data mining. The frameworks and platforms are compared among each other based on similarity, data mining task support, parallelism, distribution, streaming data processing support, fault tolerance, security, memory types, storage systems, and others. Similarly, the algorithms are grouped on the basis of parallelism type, scalability, streaming data mining support, and types of data managed. We have also provided taxonomies on the basis of data mining techniques such as clustering, classification, and association rule mining. We also have attempted to discuss and identify the major applications of cloud data mining. The various taxonomies for cloud data mining frameworks, platforms, and algorithms have been identified. This article aims at gaining better insight into the present research realm and directing the future research toward efficient cloud data mining in future cloud systems.
引用
收藏
页数:62
相关论文
共 50 条
  • [31] Metaheuristics for data mining: Survey and opportunities for big data
    Dhaenens, Clarisse
    Jourdan, Laetitia
    4OR-A QUARTERLY JOURNAL OF OPERATIONS RESEARCH, 2019, 17 (02): : 115 - 139
  • [32] Review of Algorithms for Data Mining
    Bai, Huanchen
    Liu, Xiaojun
    PROCEEDINGS OF THE 2016 6TH INTERNATIONAL CONFERENCE ON MANAGEMENT, EDUCATION, INFORMATION AND CONTROL (MEICI 2016), 2016, 135 : 383 - 386
  • [33] Data Mining and Text Mining - A Survey
    Suresh, R.
    Harshni, S. R.
    2017 INTERNATIONAL CONFERENCE ON COMPUTATION OF POWER, ENERGY INFORMATION AND COMMUNICATION (ICCPEIC), 2017, : 412 - 419
  • [34] Cloud Computing Based Data Mining of Medical Information
    Wang, Lihua
    Zhang, Ze
    PROCEEDINGS OF THE 4TH INTERNATIONAL CONFERENCE ON INFORMATION TECHNOLOGY AND MANAGEMENT INNOVATION, 2015, 28 : 978 - 982
  • [35] Improving the Performance of Data Mining by Using Big Data in Cloud Environment
    Dahmani, Djilali
    Rahal, Sid Ahmed
    Belalem, Ghalem
    JOURNAL OF INFORMATION & KNOWLEDGE MANAGEMENT, 2016, 15 (04)
  • [36] A Survey of Mass Data Mining Based on Cloud-computing
    Hu, Tingting
    Chen, Haishan
    Huang, Lu
    Zhu, Xiaodan
    2012 INTERNATIONAL CONFERENCE ON ANTI-COUNTERFEITING, SECURITY AND IDENTIFICATION (ASID), 2012,
  • [37] A Survey of Data Mining and Deep Learning in Bioinformatics
    Lan, Kun
    Wang, Dan-tong
    Fong, Simon
    Liu, Lian-sheng
    Wong, Kelvin K. L.
    Dey, Nilanjan
    JOURNAL OF MEDICAL SYSTEMS, 2018, 42 (08)
  • [38] A Survey of Social Media, Big Data, Data Mining, and Analytics
    Oliverio, Jared
    JOURNAL OF INDUSTRIAL INTEGRATION AND MANAGEMENT-INNOVATION AND ENTREPRENEURSHIP, 2018, 3 (03)
  • [39] Big Data Analytics Using Data Mining Techniques: A Survey
    Mittal, Shweta
    Sangwan, Om Prakash
    ADVANCED INFORMATICS FOR COMPUTING RESEARCH, ICAICR 2018, PT I, 2019, 955 : 264 - 273
  • [40] Design of SPRINT Parallelization of Data Mining Algorithms Based on Cloud Computing
    Song, Lei
    Zhang, Huajie
    Feng, Dongdong
    ENGINEERING LETTERS, 2022, 30 (02) : 399 - 405