Transforming Scene Text Detection and Recognition: A Multi-Scale End-to-End Approach With Transformer Framework

被引：0

作者：

Geng, Tianyu ^{[1
]}

机构：

[1] Nanjing Tech Univ, Coll Artificial Intelligence, Coll Comp & Informat Engn, Nanjing 211816, Jiangsu, Peoples R China

来源：

IEEE ACCESS | 2024年 / 12卷

关键词：

Text recognition; text recognition; transformer; end-to-end; multi-scale;

D O I：

10.1109/ACCESS.2024.3375497

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Text is an essential means for humans to acquire information and engage in social communication. Accurate text extraction from images is crucial for various tasks in real-life scenarios and scene understanding. However, text detection and recognition in natural scenes are challenged by noise in the images, irregular distribution of text fonts, and degradation of image quality under complex acquisition conditions. These factors severely impact the accuracy of text recognition. Issues such as poor image quality, diverse text formats, and complex image backgrounds significantly affect the accuracy of the recognition, and these challenges remain urgent to be addressed in the field. To address these challenges, this paper proposes a transformer-based scene image text detection and recognition algorithm within a multi-scale end-to-end framework. Firstly, by integrating detection and recognition stages into an end-to-end framework, the process is simplified, reducing computation and errors. Subsequently, multi-scale characteristics are incorporated to effectively capture text information at various scales, enhancing recognition accuracy and robustness through feature fusion and anti-interference capability. Lastly, leveraging the transformer framework, the algorithm efficiently handles text information of different scales and positions, improving generalization ability. The self-attention mechanism, multi-layer stacking structure, and positional encoding in the transformer framework contribute to its effectiveness in processing diverse text information. Through validation, the proposed method demonstrates improved efficiency in scene text detection and recognition.

引用

页码：40582 / 40596

页数：15

共 50 条

[31] End-to-End Temporal Action Detection With Transformer
Liu, Xiaolong
Wang, Qimeng
Hu, Yao
Tang, Xu
Zhang, Shiwei
Bai, Song
Bai, Xiang
IEEE TRANSACTIONS ON IMAGE PROCESSING, 2022, 31 : 5427 - 5441
[32] End-to-end speaker identification research based on multi-scale SincNet and CGAN
Guangcun Wei
Yanna Zhang
Hang Min
Yunfei Xu
Neural Computing and Applications, 2023, 35 : 22209 - 22222
[33] End-to-end Speech-to-Punctuated-Text Recognition
Nozaki, Jumon
Kawahara, Tatsuya
Ishizuka, Kenkichi
Hashimoto, Taiichi
INTERSPEECH 2022, 2022, : 1811 - 1815
[34] End-to-end reconstruction of multi-scale holograms based on CUE-NET
Wang, Shuo
Jiang, Xianan
Liu, Xu
Dong, Zhao
Pei, Ruijing
Wang, Huaying
OPTICS COMMUNICATIONS, 2023, 530
[35] End-to-end speaker identification research based on multi-scale SincNet and CGAN
Wei, Guangcun
Zhang, Yanna
Min, Hang
Xu, Yunfei
NEURAL COMPUTING & APPLICATIONS, 2023, 35 (30) : 22209 - 22222
[36] End-to-End Video Scene Graph Generation With Temporal Propagation Transformer
Zhang, Yong
Pan, Yingwei
Yao, Ting
Huang, Rui
Mei, Tao
Chen, Chang-Wen
IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 : 1613 - 1625
[37] On-device Streaming Transformer-based End-to-End Speech Recognition
Oh, Yoo Rhee
Park, Kiyoung
INTERSPEECH 2021, 2021, : 967 - 968
[38] Transformer Based End-to-End Mispronunciation Detection and Diagnosis
Wu, Minglin
Li, Kun
Leung, Wai-Kim
Meng, Helen
INTERSPEECH 2021, 2021, : 3954 - 3958
[39] Spatial–temporal transformer for end-to-end sign language recognition
Zhenchao Cui
Wenbo Zhang
Zhaoxin Li
Zhaoqi Wang
Complex & Intelligent Systems, 2023, 9 : 4645 - 4656
[40] Semantic Mask for Transformer based End-to-End Speech Recognition
Wang, Chengyi
Wu, Yu
Du, Yujiao
Li, Jinyu
Liu, Shujie
Lu, Liang
Ren, Shuo
Ye, Guoli
Zhao, Sheng
Zhou, Ming
INTERSPEECH 2020, 2020, : 971 - 975

← 1 2 3 4 5 →