High-Order Interaction Learning for Image Captioning

被引:68
|
作者
Wang, Yanhui [1 ]
Xu, Ning [1 ]
Liu, An-An [1 ]
Li, Wenhui [1 ]
Zhang, Yongdong [2 ]
机构
[1] Tianjin Univ, Sch Elect & Informat Engn, Tianjin 300072, Peoples R China
[2] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230052, Peoples R China
基金
中国国家自然科学基金;
关键词
Visualization; Semantics; Feature extraction; Decoding; Task analysis; Ions; Encoding; Image captioning; high-order interaction; encoder-decoder framework;
D O I
10.1109/TCSVT.2021.3121062
中图分类号
TM [电工技术]; TN [电子技术、通信技术];
学科分类号
0808 ; 0809 ;
摘要
Image captioning aims at understanding various semantic concepts (e.g., objects and relationships) from an image and integrating them in a sentence-level description. Hence, it is necessary to learn the interaction among these concepts. If we define the context of the interaction to be involved in the subject-predicate-object triplet, most current methods only focus on the single triplet for the first-order interaction to generate sentences. Intuitively, we humans are able to perceive the high-order interaction among concepts from two or more triplets to describe an image. For example, when we see the triplets man-cutting-sandwich and man-with-knife, it is natural to integrate and predict the sentence man cutting sandwich with knife. This depends on the high-order interaction between cutting and knife in different triplets. Therefore, exploiting high-order interaction is expected to benefit image captioning and focus on reasoning. In this paper, we introduce the novel high-order interaction learning method over detected objects and relationships for image captioning under the umbrella of the encoder-decoder framework. We first extract a set of object and relationship features in an image. During the encoding stage, the interactive refining network is proposed to learn high-order representations by modeling intra- and inter-object feature interaction in the self-attention fashion. During the decoding stage, the interactive fusion network is proposed to integrate object and relationship information by strengthening their high-order interaction based on language context for sentence generation. In this way, we learn the object-relationship dependencies in different stages, which can provide abundant cues for both visual understanding and caption generation. Extensive experiments show that the proposed method can achieve competitive performances against the state-of-the-art methods on MSCOCO dataset. Additional ablation studies further validate its effectiveness.
引用
收藏
页码:4417 / 4430
页数:14
相关论文
共 50 条
  • [1] Region-Aware Image Captioning via Interaction Learning
    Liu, An-An
    Zhai, Yingchen
    Xu, Ning
    Nie, Weizhi
    Li, Wenhui
    Zhang, Yongdong
    IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2022, 32 (06) : 3685 - 3696
  • [2] Double-Stream Position Learning Transformer Network for Image Captioning
    Jiang, Weitao
    Zhou, Wei
    Hu, Haifeng
    IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2022, 32 (11) : 7706 - 7718
  • [3] Intertemporal Interaction and Symmetric Difference Learning for Remote Sensing Image Change Captioning
    Li, Yunpeng
    Zhang, Xiangrong
    Cheng, Xina
    Chen, Puhua
    Jiao, Licheng
    IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 2024, 62
  • [4] High-Order and Interactive Perceptual Feature Learning for Medical Image Retargeting
    Ma, Mingjuan
    Zhang, Yuehong
    IEEE ACCESS, 2025, 13 : 55358 - 55369
  • [5] WordSentence Framework for Remote Sensing Image Captioning
    Wang, Qi
    Huang, Wei
    Zhang, Xueting
    Li, Xuelong
    IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 2021, 59 (12): : 10532 - 10543
  • [6] Unpaired Image Captioning With semantic-Constrained Self-Learning
    Ben, Huixia
    Pan, Yingwei
    Li, Yehao
    Yao, Ting
    Hong, Richang
    Wang, Meng
    Mei, Tao
    IEEE TRANSACTIONS ON MULTIMEDIA, 2022, 24 : 904 - 916
  • [7] Semantic-Guided Selective Representation for Image Captioning
    Li, Yinan
    Ma, Yiwei
    Zhou, Yiyi
    Yu, Xiao
    IEEE ACCESS, 2023, 11 : 14500 - 14510
  • [8] Multi-Gate Attention Network for Image Captioning
    Jiang, Weitao
    Li, Xiying
    Hu, Haifeng
    Lu, Qiang
    Liu, Bohong
    IEEE ACCESS, 2021, 9 : 69700 - 69709
  • [9] Dual Attention on Pyramid Feature Maps for Image Captioning
    Yu, Litao
    Zhang, Jian
    Wu, Qiang
    IEEE TRANSACTIONS ON MULTIMEDIA, 2022, 24 : 1775 - 1786
  • [10] Image-Relevant Entities Knowledge-Aware News Image Captioning
    Ajankar, Sonali
    Dutta, Tanima
    IEEE MULTIMEDIA, 2024, 31 (01) : 88 - 98