NOVEL APPROACH FOR BIG DATA CLASSIFICATION BASED ON HYBRID PARALLEL DIMENSIONALITY REDUCTION USING SPARK CLUSTER

被引：1

作者：

Ali, Ahmed Hussein ^{[1
]}

Abdullah, Mahmood Zaki ^{[2
]}

机构：

[1] Informat Inst Postgrad Studies, ICCI, Baghdad, Iraq

[2] Mustansiriyah Univ, Coll Engn, Baghdad, Iraq

来源：

COMPUTER SCIENCE-AGH | 2019年 / 20卷 / 04期

关键词：

big data; dimensionality reduction; parallel processing; Spark; PCA; LDA; ALGORITHM;

D O I：

10.7494/csci.2019.20.4.3307

中图分类号：

TP301 [理论、方法];

学科分类号：

081202 ;

摘要：

The big data concept has elicited studies on how to accurately and efficiently extract valuable information from a huge dataset. The major problem during big data mining is data dimensionality, which is due to the large number of dimensions in such datasets. This major consequence of high data dimensionality is that it affects the accuracy of machine learning (ML) classifiers; it also results in the wasting of time due to the presence of several redundant features in a dataset. This problem can be possibly solved using a fast feature reduction method. Hence, this study presents a fast HP-PL that is a new hybrid parallel feature reduction framework that utilizes spark to facilitate feature reduction on shared/distributed-memory clusters. An evaluation of the proposed HP-PL on the CICIDS2017 dataset showed the algorithm to be significantly faster than the conventional feature reduction techniques. The proposed technique required ?1 minute to select 4 dataset features from over 79 features and 3,000,000 samples on a 3-node cluster (a total of 21 cores). For the comparative algorithm, more than two hours was required to achieve the same feat. In the proposed system, Hadoop's distributed file system (HDFS) was used to achieve distributed storage, while Apache Spark was used as the computing engine. The model development was based on a parallel model with full consideration of the high performance and throughput of distributed computing. Conclusively, the proposed HP-PL method can achieve good accuracy with less memory and time compared to the conventional methods of feature reduction. This tool can be publicly accessed at https://github.com/ahmed/Fast-HP-PL.

引用

页码：413 / 431

页数：19

共 30 条

[1]

Agarwal S, 2017, PROCEEDINGS OF 2017 11TH INTERNATIONAL CONFERENCE ON INTELLIGENT SYSTEMS AND CONTROL (ISCO 2017), P255, DOI 10.1109/ISCO.2017.7855992

[2] An Empirical Study for PCA- and LDA-Based Feature Reduction for Gas Identification [J].

Akbar, Muhammad Ali ;

Ali, Amine Ait Si ;

Amira, Abbes ;

Bensaali, Faycal ;

Benammar, Mohieddine ;

Hassan, Muhammad ;

Bermak, Amine .

IEEE SENSORS JOURNAL, 2016, 16 (14) :5734-5746

[3] Recent trends in distributed online stream processing platform for big data: Survey [J].

Ali, Ahmed Hussein ;

Abdullah, Mahmood Zaki .

2018 1ST ANNUAL INTERNATIONAL CONFERENCE ON INFORMATION AND SCIENCES (AICIS 2018), 2018, :140-145

[4]

[Anonymous], 2009, ACM SIGKDD explorations newsletter, DOI 10.1145/1656274.1656278

[5]

[Anonymous], 2018 IEEE INT ULTR S

[6] A Parallel MapReduce Algorithm to Efficiently Support Itemset Mining on High Dimensional Data [J].

Apiletti, Daniele ;

Baralis, Elena ;

Cerquitelli, Tania ;

Garza, Paolo ;

Pulvirenti, Fabio ;

Michiardi, Pietro .

BIG DATA RESEARCH, 2017, 10 :53-69

[7] A Parallel Random Forest Algorithm for Big Data in a Spark Cloud Computing Environment [J].

Chen, Jianguo ;

Li, Kenli ;

Tang, Zhuo ;

Bilal, Kashif ;

Yu, Shui ;

Weng, Chuliang ;

Li, Keqin .

IEEE TRANSACTIONS ON PARALLEL AND DISTRIBUTED SYSTEMS, 2017, 28 (04) :919-933

[8]

Dahiya Priyanka, 2018, Procedia Computer Science, V132, P253, DOI 10.1016/j.procs.2018.05.169

[9]

Eleyan A, 2006, LECT NOTES COMPUT SC, V4105, P199

[10] Feature Selection Based on Hybridization of Genetic Algorithm and Particle Swarm Optimization [J].

Ghamisi, Pedram ;

Benediktsson, Jon Atli .

IEEE GEOSCIENCE AND REMOTE SENSING LETTERS, 2015, 12 (02) :309-313

← 1 2 3 →