NumCap: A Number-controlled Multi-caption Image Captioning Network

被引：11

作者：

Abdussalam, Amr ^{[1
]}

Ye, Zhongfu ^{[1
]}

Hawbani, Ammar ^{[2
]}

Al-Qatf, Majjed ^{[2
]}

Khan, Rashid ^{[1
]}

机构：

[1] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230027, Peoples R China

[2] Univ Sci & Technol China, Sch Comp Sci & Technol, Hefei 230027, Peoples R China

来源：

ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS | 2023年 / 19卷 / 04期

关键词：

Numbers incorporation strategy; encoder-decoder framework; image captioning; order-embedding; ATTENTION; GENERATION;

D O I：

10.1145/3576927

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Image captioning is a promising task that attracted researchers in the last few years. Existing image captioning models are primarily trained to generate one caption per image. However, an image may contain rich contents, and one caption cannot express its full details. A better solution is to describe an image with multiple captions, with each caption focusing on a specific aspect of the image. In this regard, we introduce a new number-based image captioning model that describes an image with multiple sentences. An image is annotated with multiple ground-truth captions; thus, we assign an external number to each caption to distinguish its order. Given an image-number pair as input, we could achieve different captions for the same image under different numbers. First, a number is attached to the image features to form an image-number vector (INV). Then, this vector and the corresponding caption are embedded using the order-embedding approach. Afterward, the INV's embedding is fed to a language model to generate the caption. To show the efficiency of the numbers incorporation strategy, we conduct extensive experiments using MS-COCO, Flickr30K, and Flickr8K datasets. The proposed model attains 24.1 in METEOR on MS-COCO. The achieved results demonstrate that our method is competitive with a range of state-of-the-art models and validate its ability to produce different descriptions under different given numbers.

引用

页数：24

共 61 条

[1] Image Captioning With Novel Topics Guidance and Retrieval-Based Topics Re-Weighting
Al-Qatf, Majjed
Wang, Xingfu
Hawbani, Ammar
Abdussalam, Amr
Alsamhi, Saeed Hammod
[J]. IEEE TRANSACTIONS ON MULTIMEDIA, 2023, 25 : 5984 - 5999
[2] Automatic Image and Video Caption Generation With Deep Learning: A Concise Review and Algorithmic Overlap
Amirian, Soheyla
Rasheed, Khaled
Taha, Thiab R.
Arabnia, Hamid R.
[J]. IEEE ACCESS, 2020, 8 (08): : 218386 - 218400
[3] Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Anderson, Peter
He, Xiaodong
Buehler, Chris
Teney, Damien
Johnson, Mark
Gould, Stephen
Zhang, Lei
[J]. 2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, : 6077 - 6086
[4] Image Captioning Model Using Part-of-Speech Guidance Module for Description With Diverse Vocabulary
Bae, Ju-Won
Lee, Soo-Hwan
Kim, Won-Yeol
Seong, Ju-Hyeon
Seo, Dong-Hoan
[J]. IEEE ACCESS, 2022, 10 : 45219 - 45229
[5] Social tag relevance learning via ranking-oriented neighbor voting
Cui, Chaoran
Shen, Jialie
Ma, Jun
Lian, Tao
[J]. MULTIMEDIA TOOLS AND APPLICATIONS, 2017, 76 (06) : 8831 - 8857
[6] Towards Diverse and Natural Image Descriptions via a Conditional GAN
Dai, Bo
Fidler, Sanja
Urtasun, Raquel
Lin, Dahua
[J]. 2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2017, : 2989 - 2998
[7] Topic-Based Image Caption Generation
Dash, Sandeep Kumar
Acharya, Shantanu
Pakray, Partha
Das, Ranjita
Gelbukh, Alexander
[J]. ARABIAN JOURNAL FOR SCIENCE AND ENGINEERING, 2020, 45 (04) : 3025 - 3034
[8] Generalized Zero-Shot Cross-Modal Retrieval
Dutta, Titir
Biswas, Soma
[J]. IEEE TRANSACTIONS ON IMAGE PROCESSING, 2019, 28 (12) : 5953 - 5962
[9] Fu JL, 2017, APSIPA TRANS SIGNAL, V6, DOI 10.1017/ATSIP.2017.12
[10] Hierarchical LSTMs with Adaptive Attention for Visual Captioning
Gao, Lianli
Li, Xiangpeng
Song, Jingkuan
Shen, Heng Tao
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2020, 42 (05) : 1112 - 1131

← 1 2 3 4 5 6 7 →