Interactive definition and tuning of One-Class classifiers for Document Image Classification

被引:1
作者
Girard, Nathalie [1 ]
Trullo, Roger [1 ]
Barrat, Sabine [1 ]
Ragot, Nicolas [1 ]
Ramel, Jean-Yves [1 ]
机构
[1] Univ Francois Rabelais Tours, Lab Informat, Tours, France
来源
PROCEEDINGS OF 12TH IAPR WORKSHOP ON DOCUMENT ANALYSIS SYSTEMS, (DAS 2016) | 2016年
关键词
SELECTION; SUPPORT;
D O I
10.1109/DAS.2016.46
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
With mass of data, document image classification systems have to face new trends like being able to process heterogeneous data streams efficiently. Generally, when processing data streams, few knowledge is available about the content of the possible streams. Furthermore, as getting labelled data is costly, the classification model has to be learned from few available labelled examples. To handle such specific context, we think that combining one-class classifiers could be a very interesting alternative to quickly define and tune classification systems dedicated to different document streams. The main interest of one-class classifiers is that no interdependence occurs between each classifier model allowing easy removal, addition or modification of classes of documents. Such reconfiguration will not have any impact on the other classifiers. It is also noticeable that each classifier can use a different set of features compared to the other to handle the same class or even different classes. In return, as only one class is well-specified during the learning step, one-class classifiers have to be defined carefully to obtain good performances. It is more difficult to select the representative training examples and the discriminative features with only positive examples. To overcome these difficulties, we have defined a complete framework offering different methods that can help a system designer to define and tune one-class classifier models. The aims are to make easier the selection of good training examples and of suitable features depending on the class to recognize into the document stream. For that purpose, the proposed methods compute different measures to evaluate the relevance of the available features and training examples. Moreover, a visualization of the decision space according to selected examples and features is proposed to help such a choice and, an automatic tuning is proposed for the parameters of the models according to the class to recognize when a validation stream is available. The pertinence of the proposed framework is illustrated on two different use cases (a real data stream and a public data set).
引用
收藏
页码:358 / 363
页数:6
相关论文
共 22 条
  • [1] [Anonymous], 2001, Ph.D. Thesis
  • [2] Bache K., 2013, UCI Machine Learning Repository
  • [3] Selection of relevant features and examples in machine learning
    Blum, AL
    Langley, P
    [J]. ARTIFICIAL INTELLIGENCE, 1997, 97 (1-2) : 245 - 271
  • [4] Byun Yungcheol., 2000, Proceedings of the 2000 ACM Symposium on Applied Computing, V1, P1
  • [5] Diligenti M, 2003, IEEE T PATTERN ANAL, V25, P519, DOI 10.1109/TPAMI.2003.1190578
  • [6] Eglin V, 2003, PROC INT CONF DOC, P1208
  • [7] Machine learning for intelligent processing of printed documents
    Esposito, F
    Malerba, D
    Lisi, FA
    [J]. JOURNAL OF INTELLIGENT INFORMATION SYSTEMS, 2000, 14 (2-3) : 175 - 198
  • [8] Guyon I, 2003, J MACH LEARN RES, V3, P1157, DOI DOI 10.1162/153244303322753616
  • [9] A New Feature Selection Method for One-Class Classification Problems
    Jeong, Young-Seon
    Kang, In-Ho
    Jeong, Myong-Kee
    Kong, Dongjoon
    [J]. IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS PART C-APPLICATIONS AND REVIEWS, 2012, 42 (06): : 1500 - 1509
  • [10] One-class classification with Gaussian processes
    Kemmler, Michael
    Rodner, Erik
    Wacker, Esther-Sabrina
    Denzler, Joachim
    [J]. PATTERN RECOGNITION, 2013, 46 (12) : 3507 - 3518