Transforming Scene Text Detection and Recognition: A Multi-Scale End-to-End Approach With Transformer Framework

被引:0
作者
Geng, Tianyu [1 ]
机构
[1] Nanjing Tech Univ, Coll Artificial Intelligence, Coll Comp & Informat Engn, Nanjing 211816, Jiangsu, Peoples R China
关键词
Text recognition; text recognition; transformer; end-to-end; multi-scale;
D O I
10.1109/ACCESS.2024.3375497
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Text is an essential means for humans to acquire information and engage in social communication. Accurate text extraction from images is crucial for various tasks in real-life scenarios and scene understanding. However, text detection and recognition in natural scenes are challenged by noise in the images, irregular distribution of text fonts, and degradation of image quality under complex acquisition conditions. These factors severely impact the accuracy of text recognition. Issues such as poor image quality, diverse text formats, and complex image backgrounds significantly affect the accuracy of the recognition, and these challenges remain urgent to be addressed in the field. To address these challenges, this paper proposes a transformer-based scene image text detection and recognition algorithm within a multi-scale end-to-end framework. Firstly, by integrating detection and recognition stages into an end-to-end framework, the process is simplified, reducing computation and errors. Subsequently, multi-scale characteristics are incorporated to effectively capture text information at various scales, enhancing recognition accuracy and robustness through feature fusion and anti-interference capability. Lastly, leveraging the transformer framework, the algorithm efficiently handles text information of different scales and positions, improving generalization ability. The self-attention mechanism, multi-layer stacking structure, and positional encoding in the transformer framework contribute to its effectiveness in processing diverse text information. Through validation, the proposed method demonstrates improved efficiency in scene text detection and recognition.
引用
收藏
页码:40582 / 40596
页数:15
相关论文
共 50 条
  • [41] MULTI-SCALE SCENE TEXT DETECTION VIA RESOLUTION TRANSFORM
    Cheng, Peirui
    Wang, Weiqiang
    Cai, Yuanqiang
    2019 IEEE INTERNATIONAL CONFERENCE ON MULTIMEDIA AND EXPO (ICME), 2019, : 988 - 993
  • [42] ADATS: Adaptive RoI-Align based Transformer for End-to-End Text Spotting
    Huang, Zepeng
    Wan, Qi
    Chen, Junliang
    Zhao, Xiaodong
    Ye, Kai
    Shen, Linlin
    2023 IEEE INTERNATIONAL CONFERENCE ON MULTIMEDIA AND EXPO, ICME, 2023, : 1403 - 1408
  • [43] An Investigation of Positional Encoding in Transformer-based End-to-end Speech Recognition
    Yue, Fengpeng
    Ko, Tom
    2021 12TH INTERNATIONAL SYMPOSIUM ON CHINESE SPOKEN LANGUAGE PROCESSING (ISCSLP), 2021,
  • [44] cosGCTFormer: An end-to-end driver state recognition framework
    Huang, Jing
    Liu, Tingnan
    Hu, Lin
    EXPERT SYSTEMS WITH APPLICATIONS, 2025, 261
  • [45] End-to-End Chinese Image Text Recognition with Attention Model
    Sheng, Fenfen
    Zhai, Chuanlei
    Chen, Zhineng
    Xu, Bo
    NEURAL INFORMATION PROCESSING (ICONIP 2017), PT III, 2017, 10636 : 180 - 189
  • [46] Multitask Training with Text Data for End-to-End Speech Recognition
    Wang, Peidong
    Sainath, Tara N.
    Weiss, Ron J.
    INTERSPEECH 2021, 2021, : 2566 - 2570
  • [47] On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
    Li, Jinyu
    Wu, Yu
    Gaur, Yashesh
    Wang, Chengyi
    Zhao, Rui
    Liu, Shujie
    INTERSPEECH 2020, 2020, : 1 - 5
  • [48] A novel pipeline framework for multi oriented scene text image detection and recognition
    Naiemi, Fatemeh
    Ghods, Vahid
    Khalesi, Hassan
    EXPERT SYSTEMS WITH APPLICATIONS, 2021, 170
  • [49] RESC: REfine the SCore with adaptive transformer head for end-to-end object detection
    Wang, Honglie
    Jiang, Rong
    Xu, Jian
    Sun, Shouqian
    NEURAL COMPUTING & APPLICATIONS, 2022, 34 (14) : 12017 - 12028
  • [50] LIGHTWEIGHT AND EFFICIENT END-TO-END SPEECH RECOGNITION USING LOW-RANK TRANSFORMER
    Winata, Genta Indra
    Cahyawijaya, Samuel
    Lin, Zhaojiang
    Liu, Zihan
    Fung, Pascale
    2020 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, 2020, : 6144 - 6148