3D Question Answering

被引:2
作者
Ye, Shuquan [1 ]
Chen, Dongdong [2 ]
Han, Songfang [3 ]
Liao, Jing [1 ]
机构
[1] City Univ Hong Kong, Kowloon Tong, Hong Kong, Peoples R China
[2] Microsoft Cloud AI, Redmond, WA 98052 USA
[3] Univ Calif San Diego, La Jolla, CA 92093 USA
关键词
Point cloud; scene understanding; LANGUAGE; VISION;
D O I
10.1109/TVCG.2022.3225327
中图分类号
TP31 [计算机软件];
学科分类号
081202 ; 0835 ;
摘要
Visual question answering (VQA) has experienced tremendous progress in recent years. However, most efforts have only focused on 2D image question-answering tasks. In this article, we extend VQA to its 3D counterpart, 3D question answering (3DQA), which can facilitate a machine's perception of 3D real-world scenarios. Unlike 2D image VQA, 3DQA takes the color point cloud as input and requires both appearance and 3D geometrical comprehension to answer the 3D-related questions. To this end, we propose a novel transformer-based 3DQA framework "3DQA-TR", which consists of two encoders to exploit the appearance and geometry information, respectively. Finally, the multi-modal information about the appearance, geometry, and linguistic question can attend to each other via a 3D-linguistic Bert to predict the target answers. To verify the effectiveness of our proposed 3DQA framework, we further develop the first 3DQA dataset "ScanQA", which builds on the ScanNet dataset and contains over 10 K question-answer pairs for 806 scenes. To the best of our knowledge, ScanQA is the first large-scale dataset with natural-language questions and free-form answers in 3D environments that is fully human-annotated. We also use several visualizations and experiments to investigate the astonishing diversity of the collected questions and the significant differences between this task from 2D VQA and 3D captioning. Extensive experiments on this dataset demonstrate the obvious superiority of our proposed 3DQA framework over state-of-the-art VQA frameworks and the effectiveness of our major designs. Our code and dataset will be made publicly available to facilitate research in this direction. The code and data are available at http://shuquanye.com/3DQA_website/.
引用
收藏
页码:1772 / 1786
页数:15
相关论文
共 50 条
  • [41] Improving Visual Question Answering by Multimodal Gate Fusion Network
    Xiang, Shenxiang
    Chen, Qiaohong
    Fang, Xian
    Guo, Menghao
    2023 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS, IJCNN, 2023,
  • [42] WeaQA: Weak Supervision via Captions for Visual Question Answering
    Banerjee, Pratyay
    Gokhale, Tejas
    Yang, Yezhou
    Baral, Chitta
    FINDINGS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, ACL-IJCNLP 2021, 2021, : 3420 - 3435
  • [43] A Better Way to Attend: Attention With Trees for Video Question Answering
    Xue, Hongyang
    Chu, Wenqing
    Zhao, Zhou
    Cai, Deng
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2018, 27 (11) : 5563 - 5574
  • [44] Generating 3D Model of Furniture from 3D Point Cloud of Room
    Osakama, Shunta
    Manabe, Yoshitsugu
    Yata, Noriko
    INTERNATIONAL WORKSHOP ON ADVANCED IMAGING TECHNOLOGY (IWAIT) 2020, 2020, 11515
  • [45] 3D Building Scene Reconstruction Based on 3D LiDAR Point Cloud
    Yang, Shih-Chi
    Fan, Yu-Cheng
    2017 IEEE INTERNATIONAL CONFERENCE ON CONSUMER ELECTRONICS - TAIWAN (ICCE-TW), 2017,
  • [46] GazeVQA: A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented Collaborations
    Ilaslan, Muhammet Furkan
    Song, Chenan
    Chen, Joya
    Gao, Difei
    Lei, Weixian
    Xu, Qianli
    Lim, Joo Hwee
    Shou, Mike Zheng
    2023 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (EMNLP 2023), 2023, : 10462 - 10479
  • [47] What is 3D Good For? A Review of Human Performance on Stereoscopic 3D Displays
    McIntire, John P.
    Havig, Paul R.
    Geiselman, Eric E.
    HEAD- AND HELMET-MOUNTED DISPLAYS XVII AND DISPLAY TECHNOLOGIES AND APPLICATIONS FOR DEFENSE, SECURITY, AND AVIONICS VI, 2012, 8383
  • [48] SMS3D: 3D Synthetic Mushroom Scenes Dataset for 3D Object Detection and Pose Estimation
    Zakeri, Abdollah
    Koirala, Bikram
    Kang, Jiming
    Balan, Venkatesh
    Zhu, Weihang
    Benhaddou, Driss
    Merchant, Fatima A.
    COMPUTERS, 2025, 14 (04)
  • [49] Streaming 3D Content
    Ponchio, Federico
    2ND WORKSHOP ON FLEXIBLE RESOURCE AND APPLICATION MANAGEMENT ON THE EDGE, FRAME 2022, 2022, : 1 - 2
  • [50] An analysis of graph convolutional networks and recent datasets for visual question answering
    Yusuf, Abdulganiyu Abdu
    Feng Chong
    Mao Xianling
    ARTIFICIAL INTELLIGENCE REVIEW, 2022, 55 (08) : 6277 - 6300