Comprehensive Analysis of Various Big Data Classification Techniques: A Challenging Overview

被引：4

作者：

Abdalla, Hemn Barzan ^{[1
]}

Abuhaija, Belal ^{[1
]}

机构：

[1] Wenzhou Kean Univ, Dept Comp Sci, Wenzhou, Peoples R China

来源：

JOURNAL OF INFORMATION & KNOWLEDGE MANAGEMENT | 2023年 / 22卷 / 01期

关键词：

Data mining; big data; semantic similarity measures; support vector machine; K-nearest neighbor; MAPREDUCE FRAMEWORK; ALGORITHM;

D O I：

10.1142/S0219649222500836

中图分类号：

G25 [图书馆学、图书馆事业]; G35 [情报学、情报工作];

学科分类号：

1205 ; 120501 ;

摘要：

Data over the internet has been increasing everyday, and automatic mining of essential information from an enormous amount of data has become a challenging task today for an organisation with a huge dataset. In recent years, the prominent technology in the domain of Information Technology (IT) is big data, which is unstructured data that solves the computational complexity of classical database systems. The data is fast and big and typically derived from multiple and independent sources. The three main challenges are data accessing, semantics, and domain knowledge for various big data utilisations and complexities raised by big data volumes. One of the major limitations is the classification of big data. This paper introduces well-defined classification methodologies employed for big data classification. This paper reviews 50 research papers based on classification methods of big data, and such methodologies are primarily categorised into six different categories, namely K-Nearest Neighbor (KNN), Support Vector Machine (SVM), Fuzzy-based method, Bayesian-based method, Random Forest, and Decision Tree. In addition, detailed analysis and discussion are carried out by considering classification techniques, dataset utilised, evaluation metrics, semantic similarity measures, and publication year. In addition, research gaps and issues for several traditional big data classification techniques are explained to expand investigators' works to provide effective big data management.

引用

页数：22

共 53 条

[51]

Vishwanath Brahmane Anilkumar, 2020, 2020 Fourth International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC), P851, DOI 10.1109/I-SMAC49090.2020.9243595

[52] A Classifier Ensemble Framework for Multimedia Big Data Classification [J].

Yan, Yilin ;

Zhu, Qiusha ;

Shyu, Mei-Ling ;

Chen, Shu-Ching .

PROCEEDINGS OF 2016 IEEE 17TH INTERNATIONAL CONFERENCE ON INFORMATION REUSE AND INTEGRATION (IEEE IRI), 2016, :615-622

[53] Fuzzy integral-based ELM ensemble for imbalanced big data classification [J].

Zhai, Junhai ;

Zhang, Sufang ;

Zhang, Mingyang ;

Liu, Xiaomeng .

SOFT COMPUTING, 2018, 22 (11) :3519-3531

← 1 2 3 4 5 6 →