Feature extraction using LR-PCA hybridization on twitter data and classification accuracy using machine learning algorithms

被引:41
作者
Murugan, N. Senthil [1 ]
Devi, G. Usha [1 ]
机构
[1] VIT Univ, Sch Informat Technol & Engn, Vellore, Tamil Nadu, India
来源
CLUSTER COMPUTING-THE JOURNAL OF NETWORKS SOFTWARE TOOLS AND APPLICATIONS | 2019年 / 22卷 / Suppl 6期
关键词
Social networks; Twitter; PCA; Logistic regression; Machine learning; SYSTEM;
D O I
10.1007/s10586-018-2158-3
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Twitter, a social blogging site which became the tremendous topic in today's environment, which made several organizations and public to develop their identity and overwhelming through this social website. But unfortunately, twitter facing great challenges due to spammers who break the reputation of the website from deliberate users to stop using it. Researchers have proposed many techniques to overcome the issues faced by the spammers. As far researchers find a new path so as the spammers develop new techniques to travel in that path. So far, many algorithms were proposed to detect the spammers and some extraction techniques have developed to increase the potential of detection rate. In this paper, the main focus is about feature extraction of our data with a hybrid approach of combining logistic regression with dimensional reduction technique using principal component analysis. Our dataset contains 17 million users' tweets with 159 features included in it. Then we are going to extract particular features from it which would be helpful for the further process of increasing the classification accuracy. For the classification process, our work extended for the process of classification of data using some machine learning techniques. From the proposed work the detection rate could be increased by using particular features for the classification process.
引用
收藏
页码:13965 / 13974
页数:10
相关论文
共 33 条
[1]  
[Anonymous], 2010, P COLL EL MESS ANT S
[2]   A Pattern-Based Approach for Sarcasm Detection on Twitter [J].
Bouazizi, Mondher ;
Otsuki , Tomoaki .
IEEE ACCESS, 2016, 4 :5477-5488
[3]   Statistical Features-Based Real-Time Detection of Drifted Twitter Spam [J].
Chen, Chao ;
Wang, Yu ;
Zhang, Jun ;
Xiang, Yang ;
Zhou, Wanlei ;
Min, Geyong .
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, 2017, 12 (04) :914-925
[4]  
Chen C, 2015, IEEE ICC, P7065, DOI 10.1109/ICC.2015.7249453
[5]   Sifting robotic from organic text: A natural language approach for detecting automation on Twitter [J].
Clark, Eric M. ;
Williams, Jake Ryland ;
Jones, Chris A. ;
Galbraith, Richard A. ;
Danforth, Christopher M. ;
Dodds, Peter Sheridan .
JOURNAL OF COMPUTATIONAL SCIENCE, 2016, 16 :1-7
[6]   Accelerated PSO Swarm Search Feature Selection for Data Stream Mining Big Data [J].
Fong, Simon ;
Wong, Raymond ;
Vasilakos, Athanasios V. .
IEEE TRANSACTIONS ON SERVICES COMPUTING, 2016, 9 (01) :33-45
[7]   HIoTPOT: Surveillance on IoT Devices against Recent Threats [J].
Gandhi, Usha Devi ;
Kumar, Priyan Malarvizhi ;
Varatharajan, R. ;
Manogaran, Gunasekaran ;
Sundarasekar, Revathi ;
Kadu, Shreyas .
WIRELESS PERSONAL COMMUNICATIONS, 2018, 103 (02) :1179-1194
[8]  
Gao D, 2017, SOC MEDIA CONTENT AN, DOI [10.1142/9789813223615_0024, DOI 10.1142/9789813223615_0024]
[9]   Feature extraction using independent components of each category [J].
Kotani, M ;
Ozawa, S .
NEURAL PROCESSING LETTERS, 2005, 22 (02) :113-124
[10]   RETRACTED: Intelligent face recognition and navigation system using neural learning for smart security in Internet of Things (Retracted Article) [J].
Kumar, Priyan Malarvizhi ;
Gandhi, Ushadevi ;
Varatharajan, R. ;
Manogaran, Gunasekaran ;
Jidhesh, R. ;
Vadivel, Thanjai .
CLUSTER COMPUTING-THE JOURNAL OF NETWORKS SOFTWARE TOOLS AND APPLICATIONS, 2019, 22 (Suppl 4) :S7733-S7744