Strokelets: A Learned Multi-Scale Mid-Level Representation for Scene Text Recognition

被引：72

作者：

Bai, Xiang ^{[1
]}

Yao, Cong ^{[1
]}

Liu, Wenyu ^{[1
]}

机构：

[1] Huazhong Univ Sci & Technol, Sch Elect Informat & Commun, Wuhan 430074, Peoples R China

来源：

IEEE TRANSACTIONS ON IMAGE PROCESSING | 2016年 / 25卷 / 06期

基金：

中国国家自然科学基金;

关键词：

Scene text recognition; scene text detection; mid-level representation; multi-scale representation; natural images; OBJECT DETECTION; DESCRIPTOR; VISION; MODEL;

D O I：

10.1109/TIP.2016.2555080

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

In this paper, we are concerned with the problem of automatic scene text recognition, which involves localizing and reading characters in natural images. We investigate this problem from the perspective of representation and propose a novel multi-scale representation, which leads to accurate, robust character identification and recognition. This representation consists of a set of mid-level primitives, termed strokelets, which capture the underlying substructures of characters at different granularities. The Strokelets possess four distinctive advantages: 1) usability: automatically learned from character level annotations; 2) robustness: insensitive to interference factors; 3) generality: applicable to variant languages; and 4) expressivity: effective at describing characters. Extensive experiments on standard benchmarks verify the advantages of the strokelets and demonstrate the effectiveness of the text recognition algorithm built upon the strokelets. Moreover, we show the method to incorporate the strokelets to improve the performance of scene text detection.

引用

页码：2789 / 2802

页数：14

共 84 条

[51]

Neumann L, 2012, PROC CVPR IEEE, P3538, DOI 10.1109/CVPR.2012.6248097

[52]

Neumann L, 2011, LECT NOTES COMPUT SC, V6494, P770, DOI 10.1007/978-3-642-19318-7_60

[53] Large-Lexicon Attribute-Consistent Text Recognition in Natural Images [J].

Novikova, Tatiana ;

Barinova, Olga ;

Kohli, Pushmeet ;

Lempitsky, Victor .

COMPUTER VISION - ECCV 2012, PT VI, 2012, 7577 :752-765

[54] A Hybrid Approach to Detect and Localize Texts in Natural Scene Images [J].

Pan, Yi-Feng ;

Hou, Xinwen ;

Liu, Cheng-Lin .

IEEE TRANSACTIONS ON IMAGE PROCESSING, 2011, 20 (03) :800-813

[55]

Remias L. V., 1996, REAL TIME IMAGE UNDE

[56]

Rodriguez-Serrano Jose A., 2013, P BMVC

[57]

SeongHun Lee, 2010, Proceedings of the 2010 20th International Conference on Pattern Recognition (ICPR 2010), P3983, DOI 10.1109/ICPR.2010.969

[58] ICDAR 2011 Robust Reading Competition Challenge 2: Reading Text in Scene Images [J].

Shahab, Asif ;

Shafait, Faisal ;

Dengel, Andreas .

11TH INTERNATIONAL CONFERENCE ON DOCUMENT ANALYSIS AND RECOGNITION (ICDAR 2011), 2011, :1491-1496

[59] Script identification in the wild via discriminative convolutional neural network [J].

Shi, Baoguang ;

Bai, Xiang ;

Yao, Cong .

PATTERN RECOGNITION, 2016, 52 :448-458

[60] Scene Text Recognition using Part-based Tree-structured Character Detection [J].

Shi, Cunzhao ;

Wang, Chunheng ;

Xiao, Baihua ;

Zhang, Yang ;

Gao, Song ;

Zhang, Zhong .

2013 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2013, :2961-2968

← 1 2 3 4 5 6 7 8 9 →