A novel improved random forest for text classification using feature ranking and optimal number of trees

被引:45
作者
Jalal, Nasir [1 ]
Mehmood, Arif [1 ]
Choi, Gyu Sang [2 ]
Ashraf, Imran [2 ]
机构
[1] Islamia Univ Bahawalpur, Dept Comp Sci & Informat Technol, Bahawalpur 63100, Pakistan
[2] Yeungnam Univ, Informat & Commun Engn, Gyongsan 38541, South Korea
关键词
Improved random forest; Text classification; Feature ranking; Decision tree optimization; Machine learning; Feature reduction;
D O I
10.1016/j.jksuci.2022.03.012
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Machine learning-based models like random forest (RF) have been widely deployed in diverse domains such as image processing, health care, and text processing, etc. during the past few years. The RF is a prominent technique for handling imbalanced data and performs significantly better than other machine learning models due to its parallel architecture. This study presents an improved random forest for text classification, called improved random forest for text classification (IRFTC), that incorporates bootstrapping and random subspace methods simultaneously. The IRFTC removes unimportant (less important) features, adds a number of trees in the forest on each iteration, and monitors the classification performance of RF. Classification accuracy is determined with respect to the number of trees which defines the optimal number of trees for IRFTC. Feature ranking is determined using the quality of the split in a tree. The proposed IRFTC is applied on four different benchmark datasets, binary and multiclass, to validate its performance in this study. Results indicate that IRFTC outperforms both the traditional RF, as well as, other machine learning models such as logistic regression, support vector machine, Naive Bayes, and decision trees.
引用
收藏
页码:2733 / 2742
页数:10
相关论文
共 39 条
[1]  
Almeida T., 2012, UCI Machine Learning Repository
[2]  
Almeida TA, 2011, DOCENG 2011: PROCEEDINGS OF THE 2011 ACM SYMPOSIUM ON DOCUMENT ENGINEERING, P259
[3]  
[Anonymous], 2006, U B C
[4]   MagIO: Magnetic Field Strength Based Indoor- Outdoor Detection with a Commercial Smartphone [J].
Ashraf, Imran ;
Hur, Soojung ;
Park, Yongwan .
MICROMACHINES, 2018, 9 (10)
[5]   Random forests [J].
Breiman, L .
MACHINE LEARNING, 2001, 45 (01) :5-32
[6]  
Campbell C, 2003, ACM SIGKDD EXPLORATI, V2, P1, DOI [DOI 10.1145/380995.380999, 10.1145/380995.380999]
[7]   An improved random forest classifier for multi-class classification [J].
Chaudhary A. ;
Kolhe S. ;
Kamal R. .
Information Processing in Agriculture, 2016, 3 (04) :215-222
[8]  
Choomchuay S., 2018, 2018 INT C ENG APPL, P1
[9]  
Council of Europe, 2018, HATE SPEECH
[10]  
Criminisi A., 2013, Decision Forests for Computer Vision and Medical Image Analysis