NumCap: A Number-controlled Multi-caption Image Captioning Network

被引:11
作者
Abdussalam, Amr [1 ]
Ye, Zhongfu [1 ]
Hawbani, Ammar [2 ]
Al-Qatf, Majjed [2 ]
Khan, Rashid [1 ]
机构
[1] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230027, Peoples R China
[2] Univ Sci & Technol China, Sch Comp Sci & Technol, Hefei 230027, Peoples R China
关键词
Numbers incorporation strategy; encoder-decoder framework; image captioning; order-embedding; ATTENTION; GENERATION;
D O I
10.1145/3576927
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Image captioning is a promising task that attracted researchers in the last few years. Existing image captioning models are primarily trained to generate one caption per image. However, an image may contain rich contents, and one caption cannot express its full details. A better solution is to describe an image with multiple captions, with each caption focusing on a specific aspect of the image. In this regard, we introduce a new number-based image captioning model that describes an image with multiple sentences. An image is annotated with multiple ground-truth captions; thus, we assign an external number to each caption to distinguish its order. Given an image-number pair as input, we could achieve different captions for the same image under different numbers. First, a number is attached to the image features to form an image-number vector (INV). Then, this vector and the corresponding caption are embedded using the order-embedding approach. Afterward, the INV's embedding is fed to a language model to generate the caption. To show the efficiency of the numbers incorporation strategy, we conduct extensive experiments using MS-COCO, Flickr30K, and Flickr8K datasets. The proposed model attains 24.1 in METEOR on MS-COCO. The achieved results demonstrate that our method is competitive with a range of state-of-the-art models and validate its ability to produce different descriptions under different given numbers.
引用
收藏
页数:24
相关论文
共 61 条
  • [1] Image Captioning With Novel Topics Guidance and Retrieval-Based Topics Re-Weighting
    Al-Qatf, Majjed
    Wang, Xingfu
    Hawbani, Ammar
    Abdussalam, Amr
    Alsamhi, Saeed Hammod
    [J]. IEEE TRANSACTIONS ON MULTIMEDIA, 2023, 25 : 5984 - 5999
  • [2] Automatic Image and Video Caption Generation With Deep Learning: A Concise Review and Algorithmic Overlap
    Amirian, Soheyla
    Rasheed, Khaled
    Taha, Thiab R.
    Arabnia, Hamid R.
    [J]. IEEE ACCESS, 2020, 8 (08): : 218386 - 218400
  • [3] Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
    Anderson, Peter
    He, Xiaodong
    Buehler, Chris
    Teney, Damien
    Johnson, Mark
    Gould, Stephen
    Zhang, Lei
    [J]. 2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, : 6077 - 6086
  • [4] Image Captioning Model Using Part-of-Speech Guidance Module for Description With Diverse Vocabulary
    Bae, Ju-Won
    Lee, Soo-Hwan
    Kim, Won-Yeol
    Seong, Ju-Hyeon
    Seo, Dong-Hoan
    [J]. IEEE ACCESS, 2022, 10 : 45219 - 45229
  • [5] Social tag relevance learning via ranking-oriented neighbor voting
    Cui, Chaoran
    Shen, Jialie
    Ma, Jun
    Lian, Tao
    [J]. MULTIMEDIA TOOLS AND APPLICATIONS, 2017, 76 (06) : 8831 - 8857
  • [6] Towards Diverse and Natural Image Descriptions via a Conditional GAN
    Dai, Bo
    Fidler, Sanja
    Urtasun, Raquel
    Lin, Dahua
    [J]. 2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2017, : 2989 - 2998
  • [7] Topic-Based Image Caption Generation
    Dash, Sandeep Kumar
    Acharya, Shantanu
    Pakray, Partha
    Das, Ranjita
    Gelbukh, Alexander
    [J]. ARABIAN JOURNAL FOR SCIENCE AND ENGINEERING, 2020, 45 (04) : 3025 - 3034
  • [8] Generalized Zero-Shot Cross-Modal Retrieval
    Dutta, Titir
    Biswas, Soma
    [J]. IEEE TRANSACTIONS ON IMAGE PROCESSING, 2019, 28 (12) : 5953 - 5962
  • [9] Fu JL, 2017, APSIPA TRANS SIGNAL, V6, DOI 10.1017/ATSIP.2017.12
  • [10] Hierarchical LSTMs with Adaptive Attention for Visual Captioning
    Gao, Lianli
    Li, Xiangpeng
    Song, Jingkuan
    Shen, Heng Tao
    [J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2020, 42 (05) : 1112 - 1131