Watch, attend and parse: An end-to-end neural network based approach to handwritten mathematical expression recognition

被引:152
作者
Zhang, Jianshu [1 ]
Du, Jun [1 ]
Zhang, Shiliang [1 ]
Liu, Dan [2 ]
Hu, Yulong [2 ]
Hu, Jinshui [2 ]
Wei, Si [2 ]
Dai, Lirong [1 ]
机构
[1] Univ Sci & Technol China, Natl Engn Lab Speech & Language Informat Proc, Hefei, Anhui, Peoples R China
[2] IFLYTEK Res, Hefei, Anhui, Peoples R China
基金
中国国家自然科学基金;
关键词
Handwritten mathematical expression; recognition; Neural network; Attention; FEATURES;
D O I
10.1016/j.patcog.2017.06.017
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Machine recognition of a handwritten mathematical expression (HME) is challenging due to the ambiguities of handwritten symbols and the two-dimensional structure of mathematical expressions. Inspired by recent work in deep learning, we present Watch, Attend and Parse (WAP), a novel end-to-end approach based on neural network that learns to recognize HMEs in a two-dimensional layout and outputs them as one-dimensional character sequences in LaTeX format. Inherently unlike traditional methods, our proposed model avoids problems that stem from symbol segmentation, and it does not require a predefined expression grammar. Meanwhile, the problems of symbol recognition and structural analysis are handled, respectively, using a watcher and a parser. We employ a convolutional neural network encoder that takes HME images as input as the watcher and employ a recurrent neural network decoder equipped with an attention mechanism as the parser to generate LaTeX sequences. Moreover, the correspondence between the input expressions and the output LaTeX sequences is learned automatically by the attention mechanism. We validate the proposed approach on a benchmark published by the CROHME international competition. Using the official training dataset, WAP significantly outperformed the state-of-the-art method with an expression recognition accuracy of 46.55% on CROHME 2014 and 44.55% on CROHME 2016. (C) 2017 Elsevier Ltd. All rights reserved.
引用
收藏
页码:196 / 206
页数:11
相关论文
共 62 条
[1]   An integrated grammar-based approach for mathematical expression recognition [J].
Alvaro, Francisco ;
Sanchez, Joan-Andreu ;
Benedi, Jose-Miguel .
PATTERN RECOGNITION, 2016, 51 :135-147
[2]   Recognition of on-line handwritten mathematical expressions using 2D stochastic context-free grammars and hidden Markov models [J].
Alvaro, Francisco ;
Sanchez, Joan-Andreu ;
Benedi, Jose-Miguel .
PATTERN RECOGNITION LETTERS, 2014, 35 :58-67
[3]  
Anderson R.H, 1967, S INT SYST EXP APPL, P436, DOI DOI 10.1145/2402536.2402585
[4]  
[Anonymous], 2016, What you get is what you see: a visual markup decompiler
[5]  
[Anonymous], 2016, IEEE T PATTERN ANAL
[6]  
[Anonymous], ARXIV14091556
[7]  
[Anonymous], 1994, DOCUMENT PREPARATION
[8]  
[Anonymous], INT C FRONT HANDWR R
[9]  
[Anonymous], ARXIV151107916
[10]  
[Anonymous], 2017, ACM, DOI [DOI 10.1145/3065386, DOI 10.2165/00129785-200404040-00005]