Malicious web domain identification using online credibility and performance data by considering the class imbalance issue

被引:23
作者
Hu, Zhongyi [1 ]
Chiong, Raymond [2 ]
Pranata, Ilung [2 ]
Bao, Yukun [3 ]
Lin, Yuqing [2 ]
机构
[1] Wuhan Univ, Sch Informat Management, Wuhan, Hubei, Peoples R China
[2] Univ Newcastle, Sch Elect Engn & Comp, Callaghan, NSW, Australia
[3] Huazhong Univ Sci & Technol, Sch Management, Wuhan, Hubei, Peoples R China
基金
中国博士后科学基金;
关键词
Particle swarm optimization; Imbalance class distribution; Malicious web domain; Synthetic minority oversampling technique; Online data; Credibility and performance; Information security; Internet users; SMOTE; CLASSIFICATION; PSO;
D O I
10.1108/IMDS-02-2018-0072
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
Purpose Malicious web domain identification is of significant importance to the security protection of internet users. With online credibility and performance data, the purpose of this paper to investigate the use of machine learning techniques for malicious web domain identification by considering the class imbalance issue (i.e. there are more benign web domains than malicious ones). Design/methodology/approach The authors propose an integrated resampling approach to handle class imbalance by combining the synthetic minority oversampling technique (SMOTE) and particle swarm optimisation (PSO), a population-based meta-heuristic algorithm. The authors use the SMOTE for oversampling and PSO for undersampling. Findings By applying eight well-known machine learning classifiers, the proposed integrated resampling approach is comprehensively examined using several imbalanced web domain data sets with different imbalance ratios. Compared to five other well-known resampling approaches, experimental results confirm that the proposed approach is highly effective. Practical implications - This study not only inspires the practical use of online credibility and performance data for identifying malicious web domains but also provides an effective resampling approach for handling the class imbalance issue in the area of malicious web domain identification. Originality/value Online credibility and performance data are applied to build malicious web domain identification models using machine learning techniques. An integrated resampling approach is proposed to address the class imbalance issue. The performance of the proposed approach is confirmed based on real-world data sets with different imbalance ratios.
引用
收藏
页码:676 / 696
页数:21
相关论文
共 70 条
[1]   Using Case-Based Reasoning for Phishing Detection [J].
Abutair, Hassan Y. A. ;
Belghith, Abdelfettah .
8TH INTERNATIONAL CONFERENCE ON AMBIENT SYSTEMS, NETWORKS AND TECHNOLOGIES (ANT-2017) AND THE 7TH INTERNATIONAL CONFERENCE ON SUSTAINABLE ENERGY INFORMATION TECHNOLOGY (SEIT 2017), 2017, 109 :281-288
[2]  
Agrawal A, 2015, 2015 7TH INTERNATIONAL JOINT CONFERENCE ON KNOWLEDGE DISCOVERY, KNOWLEDGE ENGINEERING AND KNOWLEDGE MANAGEMENT (IC3K), P226
[3]  
[Anonymous], 2010, NDSS 10
[4]  
[Anonymous], 2007, ICML
[5]  
[Anonymous], NATURE INSPIRED INFO
[6]   Heuristic nonlinear regression strategy for detecting phishing websites [J].
Babagoli, Mehdi ;
Aghababa, Mohammad Pourmahmood ;
Solouk, Vahid .
SOFT COMPUTING, 2019, 23 (12) :4315-4327
[7]   Strategies for learning in class imbalance problems [J].
Barandela, R ;
Sánchez, JS ;
García, V ;
Rangel, E .
PATTERN RECOGNITION, 2003, 36 (03) :849-851
[8]   MWMOTE-Majority Weighted Minority Oversampling Technique for Imbalanced Data Set Learning [J].
Barua, Sukarna ;
Islam, Md. Monirul ;
Yao, Xin ;
Murase, Kazuyuki .
IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2014, 26 (02) :405-425
[9]  
Bayes T., 1763, A. M. F. R. S. Philosophical Transactions, V53, P370
[10]   MAHAKIL: Diversity Based Oversampling Approach to Alleviate the Class Imbalance Issue in Software Defect Prediction [J].
Benni, Kwabena Ebo ;
Keung, Jacky ;
Phannachitta, Passakorn ;
Monden, Akito ;
Mensah, Solomon .
IEEE TRANSACTIONS ON SOFTWARE ENGINEERING, 2018, 44 (06) :534-550