CenterFormer: Center-Based Transformer for 3D Object Detection

被引:61
|
作者
Zhou, Zixiang [1 ,2 ]
Zhao, Xiangchen [1 ]
Wang, Yu [1 ]
Wang, Panqu [1 ]
Foroosh, Hassan [2 ]
机构
[1] TuSimple, San Diego, CA 92122 USA
[2] Univ Cent Florida, Computat Imaging Lab, Orlando, FL 32816 USA
来源
COMPUTER VISION, ECCV 2022, PT XXXVIII | 2022年 / 13698卷
关键词
LiDAR point cloud; 3D object detection; Transformer; Multi-frame fusion;
D O I
10.1007/978-3-031-19839-7_29
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Query-based transformer has shown great potential in constructing long-range attention in many image-domain tasks, but has rarely been considered in LiDAR-based 3D object detection due to the overwhelming size of the point cloud data. In this paper, we propose CenterFormer, a center-based transformer network for 3D object detection. CenterFormer first uses a center heatmap to select center candidates on top of a standard voxel-based point cloud encoder. It then uses the feature of the center candidate as the query embedding in the transformer. To further aggregate features from multiple frames, we design an approach to fuse features through cross-attention. Lastly, regression heads are added to predict the bounding box on the output center feature representation. Our design reduces the convergence difficulty and computational complexity of the transformer structure. The results show significant improvements over the strong baseline of anchor-free object detection networks. CenterFormer achieves state-of-the-art performance for a single model on the Waymo Open Dataset, with 73.7% mAPH on the validation set and 75.6% mAPH on the test set, significantly outperforming all previously published CNN and transformer-based methods. Our code is publicly available at https://github.com/TuSimple/centerformer
引用
收藏
页码:496 / 513
页数:18
相关论文
共 50 条
  • [31] CT3D++: Improving 3D Object Detection with Keypoint-Induced Channel-wise Transformer
    Sheng, Hualian
    Cai, Sijia
    Zhao, Na
    Deng, Bing
    Liang, Qiao
    Zhao, Min-Jian
    Ye, Jieping
    INTERNATIONAL JOURNAL OF COMPUTER VISION, 2025, : 4817 - 4836
  • [32] Improved Two-Stage 3D Object Detection Algorithm for Roadside Scenes with Enhanced PointPillars and Transformer
    Wang Liangzi
    Huang Miaohua
    Liu Ruoying
    Bi Chengcheng
    Hu Yongkang
    LASER & OPTOELECTRONICS PROGRESS, 2024, 61 (18)
  • [33] 3D Siamese Transformer Network for Single Object Tracking on Point Clouds
    Hui, Le
    Wang, Lingpeng
    Tang, Linghua
    Lan, Kaihao
    Xie, Jin
    Yang, Jian
    COMPUTER VISION - ECCV 2022, PT II, 2022, 13662 : 293 - 310
  • [34] Survey on deep learning-based 3D object detection in autonomous driving
    Liang, Zhenming
    Huang, Yingping
    TRANSACTIONS OF THE INSTITUTE OF MEASUREMENT AND CONTROL, 2023, 45 (04) : 761 - 776
  • [35] POAT-Net: Parallel Offset-Attention Assisted Transformer for 3D Object Detection for Autonomous Driving
    Wang, Jinyang
    Lin, Xiao
    Yu, Hongying
    IEEE ACCESS, 2021, 9 : 151110 - 151117
  • [36] GATR: Transformer Based on Guided Aggregation Decoder for 3D Multi-Modal Detection
    Luo, Yikai
    He, Linyuan
    Ma, Shiping
    Qi, Zisen
    Fan, Zunlin
    IEEE ROBOTICS AND AUTOMATION LETTERS, 2024, 9 (11): : 9725 - 9732
  • [37] FCOS3Dformer: enhancing monocular 3D object detection through transformer-assisted fusion of depth information
    Hao, Bingsen
    Deng, Zhaoxue
    Liu, Mingze
    Liu, Can
    International Journal of Vehicle Systems Modelling and Testing, 2024, 18 (03) : 228 - 244
  • [38] VoxT-GNN: A 3D object detection approach from point cloud based on voxel-level transformer and graph neural network
    Zheng, Qiangwen
    Wu, Sheng
    Wei, Jinghui
    INFORMATION PROCESSING & MANAGEMENT, 2025, 62 (04)
  • [39] 3D reconstruction of digital rocks based on StyleGAN and transformer
    Ting Zhang
    Wenqing Zhang
    International Journal of Coal Science & Technology, 2025, 12 (1)
  • [40] 3D FACIAL EXPRESSION GENERATOR BASED ON TRANSFORMER VAE
    Zou, Kaifeng
    Yu, Boyang
    Seo, Hyewon
    2023 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, ICIP, 2023, : 2550 - 2554