Strokelets: A Learned Multi-Scale Mid-Level Representation for Scene Text Recognition

被引:71
作者
Bai, Xiang [1 ]
Yao, Cong [1 ]
Liu, Wenyu [1 ]
机构
[1] Huazhong Univ Sci & Technol, Sch Elect Informat & Commun, Wuhan 430074, Peoples R China
基金
中国国家自然科学基金;
关键词
Scene text recognition; scene text detection; mid-level representation; multi-scale representation; natural images; OBJECT DETECTION; DESCRIPTOR; VISION; MODEL;
D O I
10.1109/TIP.2016.2555080
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
In this paper, we are concerned with the problem of automatic scene text recognition, which involves localizing and reading characters in natural images. We investigate this problem from the perspective of representation and propose a novel multi-scale representation, which leads to accurate, robust character identification and recognition. This representation consists of a set of mid-level primitives, termed strokelets, which capture the underlying substructures of characters at different granularities. The Strokelets possess four distinctive advantages: 1) usability: automatically learned from character level annotations; 2) robustness: insensitive to interference factors; 3) generality: applicable to variant languages; and 4) expressivity: effective at describing characters. Extensive experiments on standard benchmarks verify the advantages of the strokelets and demonstrate the effectiveness of the text recognition algorithm built upon the strokelets. Moreover, we show the method to incorporate the strokelets to improve the performance of scene text detection.
引用
收藏
页码:2789 / 2802
页数:14
相关论文
共 84 条
  • [1] [Anonymous], 2013, International Conference on Machine Learning
  • [2] [Anonymous], P ACCV
  • [3] 3D Shape Matching via Two Layer Coding
    Bai, Xiang
    Bai, Song
    Zhu, Zhuotun
    Latecki, Longin Jan
    [J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2015, 37 (12) : 2361 - 2373
  • [4] Vision-based target geo-location using a fixed-wing miniature air vehicle
    Barber, D. Blake
    Redding, Joshua D.
    McLain, Timothy W.
    Beard, Randal W.
    Taylor, Clark N.
    [J]. JOURNAL OF INTELLIGENT & ROBOTIC SYSTEMS, 2006, 47 (04) : 361 - 382
  • [5] PhotoOCR: Reading Text in Uncontrolled Conditions
    Bissacco, Alessandro
    Cummins, Mark
    Netzer, Yuval
    Neven, Hartmut
    [J]. 2013 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2013, : 785 - 792
  • [6] Poselets: Body Part Detectors Trained Using 3D Human Pose Annotations
    Bourdev, Lubomir
    Malik, Jitendra
    [J]. 2009 IEEE 12TH INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2009, : 1365 - 1372
  • [7] Bourdev L, 2010, LECT NOTES COMPUT SC, V6316, P168, DOI 10.1007/978-3-642-15567-3_13
  • [8] Random forests
    Breiman, L
    [J]. MACHINE LEARNING, 2001, 45 (01) : 5 - 32
  • [9] Chen H., 2011, 2011 18th IEEE International Conference on Image Processing (ICIP 2011), P2609, DOI 10.1109/ICIP.2011.6116200
  • [10] MEAN SHIFT, MODE SEEKING, AND CLUSTERING
    CHENG, YZ
    [J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 1995, 17 (08) : 790 - 799