NumCap: A Number-controlled Multi-caption Image Captioning Network

被引：16

作者：

Abdussalam, Amr ^{[1
]}

Ye, Zhongfu ^{[1
]}

Hawbani, Ammar ^{[2
]}

Al-Qatf, Majjed ^{[2
]}

Khan, Rashid ^{[1
]}

机构：

[1] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230027, Peoples R China

[2] Univ Sci & Technol China, Sch Comp Sci & Technol, Hefei 230027, Peoples R China

来源：

ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS | 2023年 / 19卷 / 04期

关键词：

Numbers incorporation strategy; encoder-decoder framework; image captioning; order-embedding; ATTENTION; GENERATION;

D O I：

10.1145/3576927

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Image captioning is a promising task that attracted researchers in the last few years. Existing image captioning models are primarily trained to generate one caption per image. However, an image may contain rich contents, and one caption cannot express its full details. A better solution is to describe an image with multiple captions, with each caption focusing on a specific aspect of the image. In this regard, we introduce a new number-based image captioning model that describes an image with multiple sentences. An image is annotated with multiple ground-truth captions; thus, we assign an external number to each caption to distinguish its order. Given an image-number pair as input, we could achieve different captions for the same image under different numbers. First, a number is attached to the image features to form an image-number vector (INV). Then, this vector and the corresponding caption are embedded using the order-embedding approach. Afterward, the INV's embedding is fed to a language model to generate the caption. To show the efficiency of the numbers incorporation strategy, we conduct extensive experiments using MS-COCO, Flickr30K, and Flickr8K datasets. The proposed model attains 24.1 in METEOR on MS-COCO. The achieved results demonstrate that our method is competitive with a range of state-of-the-art models and validate its ability to produce different descriptions under different given numbers.

引用

页数：24

共 61 条

[31] Small Data Challenges in Big Data Era: A Survey of Recent Progress on Unsupervised and Semi-Supervised Methods [J].

Qi, Guo-Jun ;

Luo, Jiebo .

IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2022, 44 (04) :2168-2187

[32]

Shen J, 2011, P 19 ACM INT C MULT, P639, DOI DOI 10.1145/2072298.2072405

[33] Accurate online video tagging via probabilistic hybrid modeling [J].

Shen, Jialie ;

Wang, Meng ;

Chua, Tat-Seng .

MULTIMEDIA SYSTEMS, 2016, 22 (01) :99-113

[34] Towards Optimizing Human Labeling for Interactive Image Tagging [J].

Tang, Jinhui ;

Chen, Qiang ;

Wang, Meng ;

Yan, Shuicheng ;

Chua, Tat-Seng ;

Jain, Ramesh .

ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS, 2013, 9 (04)

[35]

Vendrov Ivan., 2016, ICLR

[36]

Vinyals O, 2015, PROC CVPR IEEE, P3156, DOI 10.1109/CVPR.2015.7298935

[37] Image Captioning with Deep Bidirectional LSTMs and Multi-Task Learning [J].

Wang, Cheng ;

Yang, Haojin ;

Meinel, Christoph .

ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS, 2018, 14 (02)

[38]

Wang Qingzhong, 2018, P IEEECVF C COMPUTER, P1

[39] Cross-Modality Retrieval by Joint Correlation Learning [J].

Wang, Shuo ;

Guo, Dan ;

Xu, Xin ;

Zhuo, Li ;

Wang, Meng .

ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS, 2019, 15 (02)

[40] ARISTA - Image Search to Annotation on Billions of Web Photos [J].

Wang, Xin-Jing ;

Zhang, Lei ;

Liu, Ming ;

Li, Yi ;

Ma, Wei-Ying .

2010 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2010, :2987-2994

← 1 2 3 4 5 6 7 →