Category-Level Pose Estimation and Iterative Refinement for Monocular RGB-D Image

被引:0
作者
Bao, Yongtang [1 ]
Qi, Yutong [2 ]
Su, Chunjian [1 ]
Geng, Yanbing [3 ]
Li, Haojie [1 ]
机构
[1] Shandong Univ Sci & Technol, Coll Comp Sci & Engn, Qingdao, Peoples R China
[2] Univ Toronto, Dept Comp & Math Sci, Scarborough, ON, Canada
[3] North Univ China, Sch Data Sci & Technol, Taiyuan, Peoples R China
基金
中国国家自然科学基金;
关键词
Deep learning; category-level pose estimation; scene understanding; transformer; TRANSFORMER;
D O I
10.1145/3695877
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Category-level pose estimation is proposed to predict the 6D pose of objects under a specific category and has wide applications in fields such as robotics, virtual reality, and autonomous driving. With the development of VR/AR technology, pose estimation has gradually become a research hotspot in 3D scene understanding. However, most methods fail to fully utilize geometric and color information to solve intra-class shape variations, which leads to inaccurate prediction results. To solve the above problems, we propose a novel pose estimation and iterative refinement network, use an attention mechanism to fuse multi-modal information to obtain color features after a coordinate transformation, and design iterative modules to ensure the accuracy of object geometric features. Specifically, we use an encoder-decoder architecture to implicitly generate a coarse-grained initial pose and refine it through an iterative refinement module. In addition, due to the differences between rotation and position estimation, we design a multi-head pose decoder that utilizes the local geometry and global features. Finally, we design a transformer-based coordinate transformation attention module to extract pose-sensitive features from RGB images and supervise color information by correlating point cloud features in different coordinate systems. We train and test our network on the synthetic dataset CAMERA25 and the real dataset REAL275. Experimental results show that our method achieves state-of-the-art performance on multiple evaluation metrics.
引用
收藏
页数:20
相关论文
共 50 条
  • [21] Context-aware 6D pose estimation of known objects using RGB-D data
    Kumar, Ankit
    Shukla, Priya
    Kushwaha, Vandana
    Nandi, Gora Chand
    [J]. MULTIMEDIA TOOLS AND APPLICATIONS, 2023, 83 (17) : 52973 - 52987
  • [22] Visual Attention and Color Cues for 6D Pose Estimation on Occluded Scenarios Using RGB-D Data
    Vidal, Joel
    Lin, Chyi-Yeu
    Marti, Robert
    [J]. SENSORS, 2021, 21 (23)
  • [23] TWO-STREAM REFINEMENT NETWORK FOR RGB-D SALIENCY DETECTION
    Liu, Di
    Hu, Yaosi
    Zhang, Kao
    Chen, Zhenzhong
    [J]. 2019 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2019, : 3925 - 3929
  • [24] 3D hand pose and shape estimation from monocular RGB via efficient 2D cues
    Fenghao Zhang
    Lin Zhao
    Shengling Li
    Wanjuan Su
    Liman Liu
    Wenbing Tao
    [J]. Computational Visual Media, 2024, 10 : 79 - 96
  • [25] 3D hand pose and shape estimation from monocular RGB via efficient 2D cues
    Zhang, Fenghao
    Zhao, Lin
    Li, Shengling
    Su, Wanjuan
    Liu, Liman
    Tao, Wenbing
    [J]. COMPUTATIONAL VISUAL MEDIA, 2024, 10 (01): : 79 - 96
  • [26] CLGFormer: Cross-Level-Guided transformer for RGB-D semantic segmentation
    Li T.
    Zhou Q.
    Wu D.
    Sun M.
    Hu T.
    [J]. Multimedia Tools and Applications, 2025, 84 (11) : 9447 - 9469
  • [27] TSwinPose: Enhanced monocular 3D human pose estimation with JointFlow
    Li, Muyu
    Hu, Henan
    Xiong, Jingjing
    Zhao, Xudong
    Yan, Hong
    [J]. EXPERT SYSTEMS WITH APPLICATIONS, 2024, 249
  • [28] LEARNING MONOCULAR 3D HUMAN POSE ESTIMATION WITH SKELETAL INTERPOLATION
    Chen, Ziyi
    Sugimoto, Akihiro
    Lai, Shang-Hong
    [J]. 2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), 2022, : 4218 - 4222
  • [29] GrapesNet: Indian RGB & RGB-D vineyard image datasets for deep learning applications
    Barbole, Dhanashree K.
    Jadhav, Parul M.
    [J]. DATA IN BRIEF, 2023, 48
  • [30] Hand pose estimation based on regression method from monocular RGB cameras for handling occlusion
    Roumaissa, Bekiri
    Chaouki, Babahenini Mohamed
    [J]. MULTIMEDIA TOOLS AND APPLICATIONS, 2024, 83 (07) : 21497 - 21523