CenterFormer: Center-Based Transformer for 3D Object Detection

被引:61
|
作者
Zhou, Zixiang [1 ,2 ]
Zhao, Xiangchen [1 ]
Wang, Yu [1 ]
Wang, Panqu [1 ]
Foroosh, Hassan [2 ]
机构
[1] TuSimple, San Diego, CA 92122 USA
[2] Univ Cent Florida, Computat Imaging Lab, Orlando, FL 32816 USA
来源
COMPUTER VISION, ECCV 2022, PT XXXVIII | 2022年 / 13698卷
关键词
LiDAR point cloud; 3D object detection; Transformer; Multi-frame fusion;
D O I
10.1007/978-3-031-19839-7_29
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Query-based transformer has shown great potential in constructing long-range attention in many image-domain tasks, but has rarely been considered in LiDAR-based 3D object detection due to the overwhelming size of the point cloud data. In this paper, we propose CenterFormer, a center-based transformer network for 3D object detection. CenterFormer first uses a center heatmap to select center candidates on top of a standard voxel-based point cloud encoder. It then uses the feature of the center candidate as the query embedding in the transformer. To further aggregate features from multiple frames, we design an approach to fuse features through cross-attention. Lastly, regression heads are added to predict the bounding box on the output center feature representation. Our design reduces the convergence difficulty and computational complexity of the transformer structure. The results show significant improvements over the strong baseline of anchor-free object detection networks. CenterFormer achieves state-of-the-art performance for a single model on the Waymo Open Dataset, with 73.7% mAPH on the validation set and 75.6% mAPH on the test set, significantly outperforming all previously published CNN and transformer-based methods. Our code is publicly available at https://github.com/TuSimple/centerformer
引用
收藏
页码:496 / 513
页数:18
相关论文
共 50 条
  • [41] Lidar Point Cloud Guided Monocular 3D Object Detection
    Peng, Liang
    Liu, Fei
    Yu, Zhengxu
    Yan, Senbo
    Deng, Dan
    Yang, Zheng
    Liu, Haifeng
    Cai, Deng
    COMPUTER VISION - ECCV 2022, PT I, 2022, 13661 : 123 - 139
  • [42] Real-Time Multimodal 3D Object Detection with Transformers
    Liu, Hengsong
    Duan, Tongle
    WORLD ELECTRIC VEHICLE JOURNAL, 2024, 15 (07):
  • [43] UAV Object Detection Based on Joint YOLO and Transformer
    Gao, Yifan
    Ding, Rui
    Zhou, Fuhui
    Wu, Qihui
    2024 INTERNATIONAL CONFERENCE ON UBIQUITOUS COMMUNICATION, UCOM 2024, 2024, : 202 - 206
  • [44] A LiDAR-Camera Fusion 3D Object Detection Algorithm
    Liu, Leyuan
    He, Jian
    Ren, Keyan
    Xiao, Zhonghua
    Hou, Yibin
    INFORMATION, 2022, 13 (04)
  • [45] 2D and 3D object detection algorithms from images: A Survey
    Chen, Wei
    Li, Yan
    Tian, Zijian
    Zhang, Fan
    ARRAY, 2023, 19
  • [46] TFIENet: Transformer Fusion Information Enhancement Network for Multimodel 3-D Object Detection
    Cao, Feng
    Jin, Yufeng
    Tao, Chongben
    Luo, Xizhao
    Gao, Zhen
    Zhang, Zufeng
    Zheng, Sifa
    Zhu, Yuan
    IEEE TRANSACTIONS ON INSTRUMENTATION AND MEASUREMENT, 2024, 73
  • [47] Transformer-based difference fusion network for RGB-D salient object detection
    Cui, Zhi-Qiang
    Wang, Feng
    Feng, Zheng-Yong
    JOURNAL OF ELECTRONIC IMAGING, 2022, 31 (06)
  • [48] SARPNET: Shape attention regional proposal network for liDAR-based 3D object detection
    Ye, Yangyang
    Chen, Houjin
    Zhang, Chi
    Hao, Xiaoli
    Zhang, Zhaoxiang
    NEUROCOMPUTING, 2020, 379 : 53 - 63
  • [49] Stray Losses Study for a Power Transformer Based on 3D FEM
    Song, Zhanhai
    Wang, Yifang
    Mou, Shuai
    Wu, Zhe
    Zhu, Yinhui
    Xiang, Bingfu
    Zhou, Ce
    MECHANICAL AND ELECTRONICS ENGINEERING III, PTS 1-5, 2012, 130-134 : 3374 - +
  • [50] Transformer based 3D semantic segmentation of urban bicycle infrastructure
    Niedermueller, Armin
    Beeking, Moritz
    JOURNAL OF LOCATION BASED SERVICES, 2024,