3D Question Answering

被引:2
作者
Ye, Shuquan [1 ]
Chen, Dongdong [2 ]
Han, Songfang [3 ]
Liao, Jing [1 ]
机构
[1] City Univ Hong Kong, Kowloon Tong, Hong Kong, Peoples R China
[2] Microsoft Cloud AI, Redmond, WA 98052 USA
[3] Univ Calif San Diego, La Jolla, CA 92093 USA
关键词
Point cloud; scene understanding; LANGUAGE; VISION;
D O I
10.1109/TVCG.2022.3225327
中图分类号
TP31 [计算机软件];
学科分类号
081202 ; 0835 ;
摘要
Visual question answering (VQA) has experienced tremendous progress in recent years. However, most efforts have only focused on 2D image question-answering tasks. In this article, we extend VQA to its 3D counterpart, 3D question answering (3DQA), which can facilitate a machine's perception of 3D real-world scenarios. Unlike 2D image VQA, 3DQA takes the color point cloud as input and requires both appearance and 3D geometrical comprehension to answer the 3D-related questions. To this end, we propose a novel transformer-based 3DQA framework "3DQA-TR", which consists of two encoders to exploit the appearance and geometry information, respectively. Finally, the multi-modal information about the appearance, geometry, and linguistic question can attend to each other via a 3D-linguistic Bert to predict the target answers. To verify the effectiveness of our proposed 3DQA framework, we further develop the first 3DQA dataset "ScanQA", which builds on the ScanNet dataset and contains over 10 K question-answer pairs for 806 scenes. To the best of our knowledge, ScanQA is the first large-scale dataset with natural-language questions and free-form answers in 3D environments that is fully human-annotated. We also use several visualizations and experiments to investigate the astonishing diversity of the collected questions and the significant differences between this task from 2D VQA and 3D captioning. Extensive experiments on this dataset demonstrate the obvious superiority of our proposed 3DQA framework over state-of-the-art VQA frameworks and the effectiveness of our major designs. Our code and dataset will be made publicly available to facilitate research in this direction. The code and data are available at http://shuquanye.com/3DQA_website/.
引用
收藏
页码:1772 / 1786
页数:15
相关论文
共 50 条
  • [1] Linguistic issues behind visual question answering
    Bernardi, Raffaella
    Pezzelle, Sandro
    LANGUAGE AND LINGUISTICS COMPASS, 2021, 15 (06):
  • [2] Biomedical question answering: A survey
    Athenikos, Sofia J.
    Han, Hyoil
    COMPUTER METHODS AND PROGRAMS IN BIOMEDICINE, 2010, 99 (01) : 1 - 24
  • [3] Scene Graph Refinement Network for Visual Question Answering
    Qian, Tianwen
    Chen, Jingjing
    Chen, Shaoxiang
    Wu, Bo
    Jiang, Yu-Gang
    IEEE TRANSACTIONS ON MULTIMEDIA, 2023, 25 : 3950 - 3961
  • [4] Graph neural networks for visual question answering: a systematic review
    Yusuf, Abdulganiyu Abdu
    Feng, Chong
    Mao, Xianling
    Ally Duma, Ramadhani
    Abood, Mohammed Salah
    Chukkol, Abdulrahman Hamman Adama
    MULTIMEDIA TOOLS AND APPLICATIONS, 2023, 83 (18) : 55471 - 55508
  • [5] On the Voice-Activated Question Answering
    Rosso, Paolo
    Hurtado, Lluis-F.
    Segarra, Encarna
    Sanchis, Emilio
    IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS PART C-APPLICATIONS AND REVIEWS, 2012, 42 (01): : 75 - 85
  • [6] HIERARCHICAL RELATIONAL ATTENTION FOR VIDEO QUESTION ANSWERING
    Chowdhury, Muhammad Iqbal Hasan
    Kien Nguyen
    Sridharan, Sridha
    Fookes, Clinton
    2018 25TH IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2018, : 599 - 603
  • [7] Achieving Human Parity on Visual Question Answering
    Yan, Ming
    Xu, Haiyang
    Li, Chenliang
    Tian, Junfeng
    Bi, Bin
    Wang, Wei
    Xu, Xianzhe
    Zhang, Ji
    Huang, Songfang
    Huang, Fei
    Si, Luo
    Jin, Rong
    ACM TRANSACTIONS ON INFORMATION SYSTEMS, 2023, 41 (03)
  • [8] Wh - Question answering in children with intellectual disability
    Sanders, Eric J.
    Erickson, Karen A.
    JOURNAL OF COMMUNICATION DISORDERS, 2018, 76 : 79 - 90
  • [9] Interactive Question Answering Systems: Literature Review
    Biancofiore, Giovanni Maria
    Deldjoo, Yashar
    Di Noia, Tommaso
    Di Sciascio, Eugenio
    Narducci, Fedelucio
    ACM COMPUTING SURVEYS, 2024, 56 (09)
  • [10] Human-Adversarial Visual Question Answering
    Sheng, Sasha
    Singh, Amanpreet
    Goswami, Vedanuj
    Magana, Jose Alberto Lopez
    Thrush, Tristan
    Galuba, Wojciech
    Parikh, Devi
    Kiela, Douwe
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 34 (NEURIPS 2021), 2021, 34