Mining discriminative patches for script identification in natural scene images

被引:7
作者
Lu, Liqiong [1 ,2 ]
Wu, Dong [1 ]
Tang, Ziwei [2 ]
Yi, Yaohua [2 ]
Huang, Faliang [3 ]
机构
[1] Lingnan Normal Univ, Dept Informat Engn, Zhanjiang, Peoples R China
[2] Wuhan Univ, Sch Printing & Packaging, Wuhan, Peoples R China
[3] Nanning Normal Univ, Sch Comp & Informat Engn, Nanning, Peoples R China
关键词
Script identification; score CNN; attention CNN; discriminative patches; scene images; WORD;
D O I
10.3233/JIFS-200260
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
This paper focuses on script identification in natural scene images. Traditional CNNs (Convolution Neural Networks) cannot solve this problem perfectly for two reasons: one is the arbitrary aspect ratios of scene images which bring much difficulty to traditional CNNs with a fixed size image as the input. And the other is that some scripts with minor differences are easily confused because they share a subset of characters with the same shapes. We propose a novel approach combing Score CNN, Attention CNN and patches. Attention CNN is utilized to determine whether a patch is a discriminative patch and calculate the contribution weight of the discriminative patch to script identification of the whole image. Score CNN uses a discriminative patch as input and predict the score of each script type. Firstly patches with the same size are extracted from the scene images. Secondly these patches are used as inputs to Score CNN and Attention CNN to train two patch-level classifiers. Finally, the results of multiple discriminative patches extracted from the same image via the above two classifiers are fused to obtain the script type of this image. Using patches with the same size as inputs to CNN can avoid the problems caused by arbitrary aspect ratios of scene images. The trained classifiers can mine discriminative patches to accurately identify some confusing scripts. The experimental results show the good performance of our approach on four public datasets.
引用
收藏
页码:551 / 563
页数:13
相关论文
共 43 条
[21]   Classification of Breast Cancer Histopathological Images Using Discriminative Patches Screened by Generative Adversarial Networks [J].
Man, Rui ;
Yang, Ping ;
Xu, Bowen .
IEEE ACCESS, 2020, 8 :155362-155377
[22]   Few-shot learning for word-level scene text script identification [J].
Naosekpam, Veronica ;
Sahu, Nilkanta .
COMPUTATIONAL INTELLIGENCE, 2024, 40 (01)
[23]   Improving patch-based scene text script identification with ensembles of conjoined networks [J].
Gomez, Lluis ;
Nicolaou, Anguelos ;
Karatzas, Dimosthenis .
PATTERN RECOGNITION, 2017, 67 :85-96
[24]   Dimensionality Reduction and Feature Selection Methods for Script Identification on Document Images [J].
Poon, Bruce ;
Rahman, Saami ;
Amin, M. Ashraful ;
Yan, Hong .
INFORMATION TECHNOLOGY IN INDUSTRY, 2014, 2 (01) :1-5
[25]   Split-net: Dual transformer encoder with splitting scene text image for script identification [J].
Roy, Ayush ;
Palaiahnakote, Shivakumara ;
Pal, Umapada ;
Liu, Cheng-Lin .
PATTERN RECOGNITION LETTERS, 2025, 196 :100-108
[26]   Fine-Grained Language Identification in Scene Text Images [J].
Li, Yongrui ;
Wu, Shilian ;
Yu, Jun ;
Wang, Zengfu .
PROCEEDINGS OF THE 29TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2021, 2021, :4573-4581
[27]   Classification of aesthetic natural scene images using statistical and semantic features [J].
Biswas, Kunal ;
Shivakumara, Palaiahnakote ;
Pal, Umapada ;
Lu, Tong ;
Blumenstein, Michael ;
Llados, Josep .
MULTIMEDIA TOOLS AND APPLICATIONS, 2023, 82 (09) :13507-13532
[28]   Classification of aesthetic natural scene images using statistical and semantic features [J].
Kunal Biswas ;
Palaiahnakote Shivakumara ;
Umapada Pal ;
Tong Lu ;
Michael Blumenstein ;
Josep Lladós .
Multimedia Tools and Applications, 2023, 82 :13507-13532
[29]   Automatic script identification from document images using cluster-based templates [J].
Hochberg, J ;
Kelly, P ;
Thomas, T ;
Kerns, L .
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 1997, 19 (02) :176-181
[30]   Word Level Script Identification Using Convolutional Neural Network Enhancement for Scenic Images [J].
Mahajan, Shilpa ;
Rani, Rajneesh .
ACM TRANSACTIONS ON ASIAN AND LOW-RESOURCE LANGUAGE INFORMATION PROCESSING, 2022, 21 (04)