Category-Level Pose Estimation and Iterative Refinement for Monocular RGB-D Image

被引:0
|
作者
Bao, Yongtang [1 ]
Qi, Yutong [2 ]
Su, Chunjian [1 ]
Geng, Yanbing [3 ]
Li, Haojie [1 ]
机构
[1] Shandong Univ Sci & Technol, Coll Comp Sci & Engn, Qingdao, Peoples R China
[2] Univ Toronto, Dept Comp & Math Sci, Scarborough, ON, Canada
[3] North Univ China, Sch Data Sci & Technol, Taiyuan, Peoples R China
基金
中国国家自然科学基金;
关键词
Deep learning; category-level pose estimation; scene understanding; transformer; TRANSFORMER;
D O I
10.1145/3695877
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Category-level pose estimation is proposed to predict the 6D pose of objects under a specific category and has wide applications in fields such as robotics, virtual reality, and autonomous driving. With the development of VR/AR technology, pose estimation has gradually become a research hotspot in 3D scene understanding. However, most methods fail to fully utilize geometric and color information to solve intra-class shape variations, which leads to inaccurate prediction results. To solve the above problems, we propose a novel pose estimation and iterative refinement network, use an attention mechanism to fuse multi-modal information to obtain color features after a coordinate transformation, and design iterative modules to ensure the accuracy of object geometric features. Specifically, we use an encoder-decoder architecture to implicitly generate a coarse-grained initial pose and refine it through an iterative refinement module. In addition, due to the differences between rotation and position estimation, we design a multi-head pose decoder that utilizes the local geometry and global features. Finally, we design a transformer-based coordinate transformation attention module to extract pose-sensitive features from RGB images and supervise color information by correlating point cloud features in different coordinate systems. We train and test our network on the synthetic dataset CAMERA25 and the real dataset REAL275. Experimental results show that our method achieves state-of-the-art performance on multiple evaluation metrics.
引用
收藏
页数:20
相关论文
共 50 条
  • [1] RBP-Pose: Residual Bounding Box Projection for Category-Level Pose Estimation
    Zhang, Ruida
    Di, Yan
    Lou, Zhiqiang
    Manhardi, Fabian
    Tombari, Federico
    Ji, Xiangyang
    COMPUTER VISION - ECCV 2022, PT I, 2022, 13661 : 655 - 672
  • [2] 6D Gripper Pose Estimation from RGB-D Image
    Tang, Qirong
    Hu, Xue
    Chu, Zhugang
    Wu, Shun
    COMPUTER VISION SYSTEMS (ICVS 2019), 2019, 11754 : 120 - 125
  • [3] Pseudo View Representation Learning for Monocular RGB-D Human Pose and Shape Estimation
    Zhu, Armando
    Li, Jiefeng
    Lu, Cewu
    IEEE SIGNAL PROCESSING LETTERS, 2022, 29 : 712 - 716
  • [4] Generative Category-Level Shape and Pose Estimation with Semantic Primitives
    Li, Guanglin
    Li, Yifeng
    Ye, Zhichao
    Zhang, Qihang
    Kong, Tao
    Cui, Zhaopeng
    Zhang, Guofeng
    CONFERENCE ON ROBOT LEARNING, VOL 205, 2022, 205 : 1390 - 1400
  • [5] Category-Level 6-D Object Pose Estimation With Shape Deformation for Robotic Grasp Detection
    Yu, Sheng
    Zhai, Di-Hua
    Guan, Yuyin
    Xia, Yuanqing
    IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2025, 36 (01) : 1857 - 1871
  • [6] Category-Level 6D Pose Estimation Based on Deep Cross-Modal Feature Fusion
    Chunhui Tang
    Mingyang Zhang
    Yi Zhao
    Shouxue Shan
    Signal, Image and Video Processing, 2025, 19 (8)
  • [7] Deep-learning pipeline for object pose estimation from an rgb-d image
    No Y.C.
    Kim Y.
    Kim D.
    Han H.-G.
    Song Y.-K.
    Kim D.
    Journal of Institute of Control, Robotics and Systems, 2021, 27 (08) : 593 - 601
  • [8] Multiple-Hand 2D Pose Estimation From a Monocular RGB Image
    Mishra, Purnendu
    Sarawadekar, Kishor
    IEEE ACCESS, 2024, 12 : 40722 - 40735
  • [9] Category-Level Object Detection, Pose Estimation and Reconstruction from Stereo Images
    Zhang, Chuanrui
    Ling, Yonggen
    Lu, Minglei
    Qin, Minghan
    Wang, Haoqian
    COMPUTER VISION - ECCV 2024, PT XXXIV, 2025, 15092 : 332 - 349
  • [10] A RGB-D feature fusion network for occluded object 6D pose estimation
    Song, Yiwei
    Tang, Chunhui
    SIGNAL IMAGE AND VIDEO PROCESSING, 2024, 18 (8-9) : 6309 - 6319