New margin-based subsampling iterative technique in modified random forests for classification

被引:45
作者
Feng, Wei [1 ,2 ]
Dauphin, Gabriel [3 ]
Huang, Wenjiang [2 ]
Quan, Yinghui [4 ]
Liao, Wenzhi [5 ]
机构
[1] Xidian Univ, Sch Elect Engn, Xian 710071, Shaanxi, Peoples R China
[2] Chinese Acad Sci, Inst Remote Sensing & Digital Earth, Key Lab Digital Earth Sci, Beijing 100094, Peoples R China
[3] Univ Paris XIII, Inst Galilee, Lab Informat Proc & Transmiss, L2TI, Villetaneuse, France
[4] Xidian Univ, Key Lab Radar Signal Proc, Xian 710071, Shaanxi, Peoples R China
[5] Univ Ghent, Dept Telecommun & Informat Proc, IMEC, TELIN, Sint Pietersnieuwstr 41, B-9000 Ghent, Belgium
基金
中国国家自然科学基金;
关键词
Classification; Ensemble margin; Diversity; Random forests; Sub-sampling; DIVERSITY; MACHINE; IMPROVEMENT; ACCURACY;
D O I
10.1016/j.knosys.2019.07.016
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Diversity within base classifiers has been recognized as an important characteristic of an ensemble classifier. Data and feature sampling are two popular methods of increasing such diversity. This is exemplified by Random Forests (RFs), known as a very effective classifier. However real-world data remain challenging due to several issues, such as multi-class imbalance, data redundancy, and class noise. Ensemble margin theory is a proven effective way to improve the performance of classification models. It can be used to detect the most important instances and thus help ensemble classifiers to avoid the negative effects of the class noise and class imbalance. To obtain accurate classification results, this paper proposes the Ensemble-Margin Based Random Forests (EMRFs) method, which combines RFs and a new subsampling iterative technique making use of computed ensemble margin values. As for comparative analysis, the learning techniques considered are: SVM, AdaBoost, RFs and the Subsample based Random Forests (SubRFs). The SubRFs uses Out-Of-Bag (OOB) estimation to optimize the training size. The effectiveness of EMRFs is demonstrated on both balanced and imbalanced datasets. (C) 2019 Elsevier B.V. All rights reserved.
引用
收藏
页数:12
相关论文
共 56 条
[1]   Increasing diversity in random forest learning algorithm via imprecise probabilities [J].
Abellan, Joaquin ;
Mantas, Carlos J. ;
Castellano, Javier G. ;
Moral-Garcia, SerafIn .
EXPERT SYSTEMS WITH APPLICATIONS, 2018, 97 :228-243
[2]  
[Anonymous], 1996, Technical Report
[3]  
[Anonymous], THESIS
[4]  
Asuncion A., 2007, Uci machine learning repository, university of california, irvine, school of information and computer sciences
[5]   Dynamic Random Forests [J].
Bernard, Simon ;
Adam, Sebastien ;
Heutte, Laurent .
PATTERN RECOGNITION LETTERS, 2012, 33 (12) :1580-1586
[6]   Stacked ensemble combined with fuzzy matching for biomedical named entity recognition of diseases [J].
Bhasuran, Balu ;
Murugesan, Gurusamy ;
Abdulkadhar, Sabenabanu ;
Natarajan, Jeyakumar .
JOURNAL OF BIOMEDICAL INFORMATICS, 2016, 64 :1-9
[7]   SmcHD1, containing a structural-maintenance-of-chromosomes hinge domain, has a critical role in X inactivation [J].
Blewitt, Marnie E. ;
Gendrel, Anne-Valerie ;
Pang, Zhenyi ;
Sparrow, Duncan B. ;
Whitelaw, Nadia ;
Craig, Jeffrey M. ;
Apedaile, Anwyn ;
Hilton, Douglas J. ;
Dunwoodie, Sally L. ;
Brockdorff, Neil ;
Kay, Graham F. ;
Whitelaw, Emma .
NATURE GENETICS, 2008, 40 (05) :663-669
[8]   Random forests [J].
Breiman, L .
MACHINE LEARNING, 2001, 45 (01) :5-32
[9]   Random forests [J].
Breiman, L .
MACHINE LEARNING, 2001, 45 (01) :5-32
[10]   Pasting small votes for classification in large databases and on-line [J].
Breiman, L .
MACHINE LEARNING, 1999, 36 (1-2) :85-103