Automated Image Captioning Using Sparrow Search Algorithm With Improved Deep Learning Model

被引:0
作者
Arasi, Munya A. [1 ]
Alshahrani, Haya Mesfer [2 ]
Alruwais, Nuha [3 ]
Motwakel, Abdelwahed [4 ]
Ahmed, Noura Abdelaziz [5 ]
Mohamed, Abdullah [6 ]
机构
[1] King Khalid Univ, Coll Sci & Arts Rijal Almaa, Dept Comp Sci, Abha 62529, Saudi Arabia
[2] Princess Nourah Bint Abdulrahman Univ, Coll Comp & Informat Sci, Dept Informat Syst, POB 84428, Riyadh 11671, Saudi Arabia
[3] King Saud Univ, Coll Appl Studies & Community Serv, Dept Comp Sci & Engn, POB 22459, Riyadh 11495, Saudi Arabia
[4] Prince Sattam Bin Abdulaziz Univ, Coll Business Adm Hawtat Bani Tamim, Dept Management Informat Syst, Al Kharj 11942, Saudi Arabia
[5] Prince Sattam Bin Abdulaziz Univ, Dept Comp & Self Dev, Preparatory Year Deanship, Al Kharj 11942, Saudi Arabia
[6] Future Univ Egypt, Res Ctr, New Cairo 11845, Egypt
关键词
Convolutional neural networks; Visualization; Feature extraction; Convolution; Deep learning; Natural language processing; Computational modeling; Image capture; Search methods; Image captioning; deep learning; natural language processing; sparrow search algorithm; computer vision;
D O I
10.1109/ACCESS.2023.3317276
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Image captioning is a deep learning technique that intends to create and generate textual descriptions or captions for images. It integrates computer vision and natural language processing (NLP) to comprehend the visual content of an image and generate human-like descriptions. Deep learning (DL) based image captioning models can be trained on large-scale datasets, allowing them to generalize various types of images and generate captions that apply to a wide range of visual scenarios. By combining computer vision and natural language processing, DL-enabled image captioning models can understand both visual and textual information, which enables them to generate captions that not only describe the visual content but also incorporate contextual and semantic information. This study develops an Automated Image Captioning using Sparrow Search Algorithm with Improved Deep Learning (AIC-SSAIDL) technique. The major intention of the AIC-SSAIDL technique lies in the automated generation of textual captions for the input images. To accomplish this, the AIC-SSAIDL technique utilizes the MobileNetv2 model to generate feature descriptors of the input images and its hyperparameter tuning process takes place using SSA. For the image captioning process, the AIC-SSAIDL technique utilizes an attention mechanism with long short-term memory (AM-LSTM) network. Finally, the hyperparameter selection of the AM-LSTM model is performed by the fruit fly optimization (FFO) algorithm. A wide range of experiments has been conducted on benchmark data to depict the better performance of the AIC-SSAIDL method. The comprehensive result analysis highlighted the enhanced captioning results of the AIC-SSAIDL method with maximum CIDEr of 46.12, 61.89, and 137.45 on Flickr8k, Flickr30k, and MSCOCO datasets, respectively.
引用
收藏
页码:104633 / 104642
页数:10
相关论文
共 28 条
[1]   Metaheuristics Optimization with Deep Learning Enabled Automated Image Captioning System [J].
Al Duhayyim, Mesfer ;
Alazwari, Sana ;
Mengash, Hanan Abdullah ;
Marzouk, Radwa ;
Alzahrani, Jaber S. ;
Mahgoub, Hany ;
Althukair, Fahd ;
Salama, Ahmed S. .
APPLIED SCIENCES-BASEL, 2022, 12 (15)
[2]   Image captioning model using attention and object features to mimic human image understanding [J].
Al-Malla, Muhammad Abdelhadie ;
Jafar, Assef ;
Ghneim, Nada .
JOURNAL OF BIG DATA, 2022, 9 (01)
[3]   Boosting convolutional image captioning with semantic content and visual relationship [J].
Bai, Cong ;
Zheng, Anqi ;
Huang, Yuan ;
Pan, Xiang ;
Chen, Nan .
DISPLAYS, 2021, 70
[4]   Deep Learning Approaches Based on Transformer Architectures for Image Captioning Tasks [J].
Castro, Roberto ;
Pineda, Israel ;
Lim, Wansu ;
Morocho-Cayamcela, Manuel Eugenio .
IEEE ACCESS, 2022, 10 :33679-33694
[5]   Improved Framework using Rider Optimization Algorithm for Precise Image Caption Generation [J].
Chaudhari, Chaitrali Prasanna ;
Devane, Satish .
INTERNATIONAL JOURNAL OF IMAGE AND GRAPHICS, 2022, 22 (02)
[6]   SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning [J].
Chen, Long ;
Zhang, Hanwang ;
Xiao, Jun ;
Nie, Liqiang ;
Shao, Jian ;
Liu, Wei ;
Chua, Tat-Seng .
30TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2017), 2017, :6298-6306
[7]  
Chu Y., 2020, Wireless Commun. Mobile Comput., P1
[8]   Image Captioning using Hybrid LSTM-RNN with Deep Features [J].
Deorukhkar, Kalpana Prasanna ;
Ket, Satish .
SENSING AND IMAGING, 2022, 23 (01)
[9]  
Elhagry A, 2021, arXiv
[10]   Attention meets long short-term memory: A deep learning network for traffic flow forecasting [J].
Fang, Weiwei ;
Zhuo, Wenhao ;
Yan, Jingwen ;
Song, Youyi ;
Jiang, Dazhi ;
Zhou, Teng .
PHYSICA A-STATISTICAL MECHANICS AND ITS APPLICATIONS, 2022, 587