Convolutional recurrent neural networks with hidden Markov model bootstrap for scene text recognition

被引:12
作者
Wang, Fenglei [1 ]
Guo, Qiang [1 ]
Lei, Jun [1 ]
Zhang, Jun [1 ]
机构
[1] Natl Univ Def Technol, Dept Informat Syst & Management, Changsha, Hunan, Peoples R China
基金
中国国家自然科学基金;
关键词
recurrent neural nets; text detection; hidden Markov models; convolutional recurrent neural networks; scene text recognition; RNN; CNN; Gaussian mixture model-hidden Markov model; lexicon-free text; lexicon-based text; HANDWRITING RECOGNITION;
D O I
10.1049/iet-cvi.2016.0417
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Text recognition in natural scene remains a challenging problem due to the highly variable appearance in unconstrained condition. The authors develop a system that directly transcribes scene text images to text without character segmentation. They formulate the problem as sequence labelling. They build a convolutional recurrent neural network (RNN) by using deep convolutional neural networks (CNN) for modelling text appearance and RNNs for sequence dynamics. The two models are complementary in modelling capabilities and so integrated together to form the segmentation free system. They train a Gaussian mixture model-hidden Markov model to supervise the training of the CNN model. The system is data driven and needs no hand labelled training data. Their method has several appealing properties: (i) It can recognise arbitrary length text images. (ii) The recognition process does not involve sophisticated character segmentation. (iii) It is trained on scene text images with only word-level transcriptions. (iv) It can recognise both the lexicon-based or lexicon-free text. The proposed system achieves competitive performance comparison with the state of the art on several public scene text datasets, including both lexicon-based and non-lexicon ones.
引用
收藏
页码:497 / 504
页数:8
相关论文
共 50 条
[31]   Object Recognition using Cellular Simultaneous Recurrent Networks and Convolutional Neural Network [J].
Alom, Md Zahangir ;
Alam, M. ;
Taha, Tarek M. ;
Iftekharuddin, K. M. .
2017 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS (IJCNN), 2017, :2873-2880
[32]   A Residual-Attention Offline Handwritten Chinese Text Recognition Based on Fully Convolutional Neural Networks [J].
Wang, Yintong ;
Yang, Yingjie ;
Ding, Weiping ;
Li, Shuo .
IEEE ACCESS, 2021, 9 :132301-132310
[33]   Automatic Segmentation and Recognition in Body Sensor Networks Using a Hidden Markov Model [J].
Guenterberg, Eric ;
Ghasemzadeh, Hassan ;
Jafari, Roozbeh .
ACM TRANSACTIONS ON EMBEDDED COMPUTING SYSTEMS, 2012, 11
[34]   ONLINE CURSIVE SCRIPT RECOGNITION USING TIME-DELAY NEURAL NETWORKS AND HIDDEN MARKOV-MODELS [J].
SCHENKEL, M ;
GUYON, I ;
HENDERSON, D .
MACHINE VISION AND APPLICATIONS, 1995, 8 (04) :215-223
[35]   A review of Hidden Markov models and Recurrent Neural Networks for event detection and localization in biomedical signals [J].
Khalifa, Yassin ;
Mandic, Danilo ;
Sejdic, Ervin .
INFORMATION FUSION, 2021, 69 :52-72
[36]   Reading Text in the Wild with Convolutional Neural Networks [J].
Max Jaderberg ;
Karen Simonyan ;
Andrea Vedaldi ;
Andrew Zisserman .
International Journal of Computer Vision, 2016, 116 :1-20
[37]   A Novel Text Sample Selection Model for Scene Text Detection via Bootstrap Learning [J].
Kong, Jun ;
Sun, Jinhua ;
Jiang, Min ;
Hou, Jian .
KSII TRANSACTIONS ON INTERNET AND INFORMATION SYSTEMS, 2019, 13 (02) :771-789
[38]   Reading Text in the Wild with Convolutional Neural Networks [J].
Jaderberg, Max ;
Simonyan, Karen ;
Vedaldi, Andrea ;
Zisserman, Andrew .
INTERNATIONAL JOURNAL OF COMPUTER VISION, 2016, 116 (01) :1-20
[39]   DIFFUSIONSTR: DIFFUSION MODEL FOR SCENE TEXT RECOGNITION [J].
Fujitake, Masato .
2023 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, ICIP, 2023, :1585-1589
[40]   Triggered Attention Model for Scene Text Recognition [J].
Zhang, Churong ;
Ming, Yue .
ELEVENTH INTERNATIONAL CONFERENCE ON GRAPHICS AND IMAGE PROCESSING (ICGIP 2019), 2020, 11373