Comprehensive Analysis of Various Big Data Classification Techniques: A Challenging Overview

被引:4
作者
Abdalla, Hemn Barzan [1 ]
Abuhaija, Belal [1 ]
机构
[1] Wenzhou Kean Univ, Dept Comp Sci, Wenzhou, Peoples R China
关键词
Data mining; big data; semantic similarity measures; support vector machine; K-nearest neighbor; MAPREDUCE FRAMEWORK; ALGORITHM;
D O I
10.1142/S0219649222500836
中图分类号
G25 [图书馆学、图书馆事业]; G35 [情报学、情报工作];
学科分类号
1205 ; 120501 ;
摘要
Data over the internet has been increasing everyday, and automatic mining of essential information from an enormous amount of data has become a challenging task today for an organisation with a huge dataset. In recent years, the prominent technology in the domain of Information Technology (IT) is big data, which is unstructured data that solves the computational complexity of classical database systems. The data is fast and big and typically derived from multiple and independent sources. The three main challenges are data accessing, semantics, and domain knowledge for various big data utilisations and complexities raised by big data volumes. One of the major limitations is the classification of big data. This paper introduces well-defined classification methodologies employed for big data classification. This paper reviews 50 research papers based on classification methods of big data, and such methodologies are primarily categorised into six different categories, namely K-Nearest Neighbor (KNN), Support Vector Machine (SVM), Fuzzy-based method, Bayesian-based method, Random Forest, and Decision Tree. In addition, detailed analysis and discussion are carried out by considering classification techniques, dataset utilised, evaluation metrics, semantic similarity measures, and publication year. In addition, research gaps and issues for several traditional big data classification techniques are explained to expand investigators' works to provide effective big data management.
引用
收藏
页数:22
相关论文
共 53 条
[51]  
Vishwanath Brahmane Anilkumar, 2020, 2020 Fourth International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC), P851, DOI 10.1109/I-SMAC49090.2020.9243595
[52]   A Classifier Ensemble Framework for Multimedia Big Data Classification [J].
Yan, Yilin ;
Zhu, Qiusha ;
Shyu, Mei-Ling ;
Chen, Shu-Ching .
PROCEEDINGS OF 2016 IEEE 17TH INTERNATIONAL CONFERENCE ON INFORMATION REUSE AND INTEGRATION (IEEE IRI), 2016, :615-622
[53]   Fuzzy integral-based ELM ensemble for imbalanced big data classification [J].
Zhai, Junhai ;
Zhang, Sufang ;
Zhang, Mingyang ;
Liu, Xiaomeng .
SOFT COMPUTING, 2018, 22 (11) :3519-3531