Semi-supervised learning using multiple clusterings with limited labeled data

被引:40
作者
Forestier, Germain [1 ]
Wemmert, Cedric [2 ]
机构
[1] Univ Haute Alsace, MIPS, Mulhouse, France
[2] Univ Strasbourg, ICube, Strasbourg, France
关键词
Semi-supervised learning; Classification; Pattern recognition; Remote sensing; UNLABELED DATA; CLASSIFICATION; FRAMEWORK; DIVERSITY; ENSEMBLE;
D O I
10.1016/j.ins.2016.04.040
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Supervised classification consists in learning a predictive model using a set of labeled samples. It is accepted that predictive models accuracy usually increases as more labeled samples are available. Labeled samples are generally difficult to obtain as the labeling step if often performed manually. On the contrary, unlabeled samples are easily available. As the labeling task is tedious and time consuming, users generally provide a very limited number of labeled objects. However, designing approaches able to work efficiently with a very limited number of labeled samples is highly challenging. In this context, semi-supervised approaches have been proposed to leverage from both labeled and unlabeled data. In this paper, we focus on cases where the number of labeled samples is very limited. We review and formalize eight semi-supervised learning algorithms and introduce a new method that combine supervised and unsupervised learning in order to use both labeled and unlabeled data. The main idea of this method is to produce new features derived from a first step of data clustering. These features are then used to enrich the description of the input data leading to a better use of the data distribution. The efficiency of all the methods is compared on various artificial, UCI datasets, and on the classification of a very high resolution remote sensing image. The experiments reveal that our method shows good results, especially when the number of labeled sample is very limited. It also confirms that combining labeled and unlabeled data is very useful in pattern recognition. (C) 2016 Elsevier Inc. All rights reserved.
引用
收藏
页码:48 / 65
页数:18
相关论文
共 49 条
  • [1] Semi-Supervised Kernel Mean Shift Clustering
    Anand, Saket
    Mittal, Sushil
    Tuzel, Oncel
    Meer, Peter
    [J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2014, 36 (06) : 1201 - 1215
  • [2] [Anonymous], 2006, INFORM FUSION 2006 9
  • [3] [Anonymous], 2002, P 8 ACM SIGKDD INT C
  • [4] [Anonymous], 2007, P 24 INT C MACH LEAR
  • [5] [Anonymous], 1998, UCI REPOSITORY MACHI
  • [6] [Anonymous], 2000, P INT C MACH LEARN I
  • [7] Combining supervised and unsupervised models via unconstrained probabilistic embedding
    Ao, Xiang
    Luo, Ping
    Ma, Xudong
    Zhuang, Fuzhen
    He, Qing
    Shi, Zhongzhi
    Shen, Zhiyong
    [J]. INFORMATION SCIENCES, 2014, 257 : 101 - 114
  • [8] Basu S., 2002, P INT C MACH LEARN, P27
  • [9] Belkin M, 2006, J MACH LEARN RES, V7, P2399
  • [10] Blum A., 1998, Proceedings of the Eleventh Annual Conference on Computational Learning Theory, P92, DOI 10.1145/279943.279962