Learning to Answer Questions from Image Using Convolutional Neural Network

被引:0
|
作者
Ma, Lin [1 ]
Lu, Zhengdong [1 ]
Li, Hang [1 ]
机构
[1] Huawei Technol, Noahs Ark Lab, Shenzhen, Peoples R China
关键词
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
In this paper, we propose to employ the convolutional neural network (CNN) for the image question answering (QA) task. Our proposed CNN provides an end-to-end framework with convolutional architectures for learning not only the image and question representations, but also their inter-modal interactions to produce the answer. More specifically, our model consists of three CNNs: one image CNN to encode the image content, one sentence CNN to compose the words of the question, and one multimodal convolution layer to learn their joint representation for the classification in the space of candidate answer words. We demonstrate the efficacy of our proposed model on the DAQUAR and COCO-QA datasets, which are two benchmark datasets for image QA, with the performances significantly outperforming the state-of-the-art.
引用
收藏
页码:3567 / 3573
页数:7
相关论文
共 50 条
  • [1] Image Denoising using Deep Learning: Convolutional Neural Network
    Ghose, Shreyasi
    Singh, Nishi
    Singh, Prabhishek
    PROCEEDINGS OF THE CONFLUENCE 2020: 10TH INTERNATIONAL CONFERENCE ON CLOUD COMPUTING, DATA SCIENCE & ENGINEERING, 2020, : 511 - 517
  • [2] Transfer learning for Hyperspectral image classification using convolutional neural network
    Liu, Yao
    Xiao, Chenchao
    MIPPR 2019: REMOTE SENSING IMAGE PROCESSING, GEOGRAPHIC INFORMATION SYSTEMS, AND OTHER APPLICATIONS, 2020, 11432
  • [3] LEARNING AND TRANSFERRING REPRESENTATIONS FOR IMAGE STEGANALYSIS USING CONVOLUTIONAL NEURAL NETWORK
    Qian, Yinlong
    Dong, Jing
    Wang, Wei
    Tan, Tieniu
    2016 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2016, : 2752 - 2756
  • [4] Scene Recognition from Image Using Convolutional Neural Network
    Masood, Sarfaraz
    Ahsan, Umer
    Munawwar, Fatima
    Rizvi, Danish Raza
    Ahmed, Mumtaz
    INTERNATIONAL CONFERENCE ON COMPUTATIONAL INTELLIGENCE AND DATA SCIENCE, 2020, 167 : 1005 - 1012
  • [5] Neural architectures for learning to answer questions
    Monner, Derek
    Reggia, James A.
    BIOLOGICALLY INSPIRED COGNITIVE ARCHITECTURES, 2012, 2 : 37 - 53
  • [6] Image Synthesis using Convolutional Neural Network
    Bhat, Ganesh
    Dharwadkar, Shrikant
    Reddy, N. V. Subba
    Shivaprasad, G.
    2017 2ND IEEE INTERNATIONAL CONFERENCE ON RECENT TRENDS IN ELECTRONICS, INFORMATION & COMMUNICATION TECHNOLOGY (RTEICT), 2017, : 689 - 691
  • [7] Image enhancement using convolutional neural network
    Zhou, Abel
    Tan, Qi
    Davidson, Rob
    2020 INTERNATIONAL CONFERENCE ON IMAGE, VIDEO PROCESSING AND ARTIFICIAL INTELLIGENCE, 2020, 11584
  • [8] Image Denoising using Convolutional Neural Network
    Mehmood, Asif
    PATTERN RECOGNITION AND TRACKING XXXI, 2020, 11400
  • [9] Medical image denoising using convolutional neural network: a residual learning approach
    Jifara, Worku
    Jiang, Feng
    Rho, Seungmin
    Cheng, Maowei
    Liu, Shaohui
    JOURNAL OF SUPERCOMPUTING, 2019, 75 (02): : 704 - 718
  • [10] A SHUFFLED DILATED CONVOLUTIONAL NEURAL NETWORK FOR HYPERSPECTRAL IMAGE USING TRANSFER LEARNING
    Chhapariya, Koushikey
    Buddhiraju, Krishna Mohan
    Kumar, Anil
    IGARSS 2023 - 2023 IEEE INTERNATIONAL GEOSCIENCE AND REMOTE SENSING SYMPOSIUM, 2023, : 7547 - 7550