BAFusion: Bidirectional Attention Fusion for 3D Object Detection Based on LiDAR and Camera

被引:4
|
作者
Liu, Min [1 ]
Jia, Yuanjun [2 ]
Lyu, Youhao [1 ]
Dong, Qi [2 ]
Yang, Yanyu [2 ]
机构
[1] Univ Sci & Technol China, Inst Adv Technol, Hefei 230088, Peoples R China
[2] China Acad Elect & Informat Technol, Beijing 100041, Peoples R China
关键词
3D object detection; LiDAR-camera fusion; cross attention;
D O I
10.3390/s24144718
中图分类号
O65 [分析化学];
学科分类号
070302 ; 081704 ;
摘要
3D object detection is a challenging and promising task for autonomous driving and robotics, benefiting significantly from multi-sensor fusion, such as LiDAR and cameras. Conventional methods for sensor fusion rely on a projection matrix to align the features from LiDAR and cameras. However, these methods often suffer from inadequate flexibility and robustness, leading to lower alignment accuracy under complex environmental conditions. Addressing these challenges, in this paper, we propose a novel Bidirectional Attention Fusion module, named BAFusion, which effectively fuses the information from LiDAR and cameras using cross-attention. Unlike the conventional methods, our BAFusion module can adaptively learn the cross-modal attention weights, making the approach more flexible and robust. Moreover, drawing inspiration from advanced attention optimization techniques in 2D vision, we developed the Cross Focused Linear Attention Fusion Layer (CFLAF Layer) and integrated it into our BAFusion pipeline. This layer optimizes the computational complexity of attention mechanisms and facilitates advanced interactions between image and point cloud data, showcasing a novel approach to addressing the challenges of cross-modal attention calculations. We evaluated our method on the KITTI dataset using various baseline networks, such as PointPillars, SECOND, and Part-A2, and demonstrated consistent improvements in 3D object detection performance over these baselines, especially for smaller objects like cyclists and pedestrians. Our approach achieves competitive results on the KITTI benchmark.
引用
收藏
页数:26
相关论文
共 33 条
  • [1] A LiDAR-Camera Fusion 3D Object Detection Algorithm
    Liu, Leyuan
    He, Jian
    Ren, Keyan
    Xiao, Zhonghua
    Hou, Yibin
    INFORMATION, 2022, 13 (04)
  • [2] LiDAR-Camera Fusion in Perspective View for 3D Object Detection in Surface Mine
    Ai, Yunfeng
    Yang, Xue
    Song, Ruiqi
    Cui, Chenglin
    Li, Xinqing
    Cheng, Qi
    Tian, Bin
    Chen, Long
    IEEE TRANSACTIONS ON INTELLIGENT VEHICLES, 2024, 9 (02): : 3721 - 3730
  • [3] FusionRCNN: LiDAR-Camera Fusion for Two-Stage 3D Object Detection
    Xu, Xinli
    Dong, Shaocong
    Xu, Tingfa
    Ding, Lihe
    Wang, Jie
    Jiang, Peng
    Song, Liqiang
    Li, Jianan
    REMOTE SENSING, 2023, 15 (07)
  • [4] PLC-Fusion: Perspective-Based Hierarchical and Deep LiDAR Camera Fusion for 3D Object Detection in Autonomous Vehicles
    Mushtaq, Husnain
    Deng, Xiaoheng
    Azhar, Fizza
    Ali, Mubashir
    Sherazi, Hafiz Husnain Raza
    INFORMATION, 2024, 15 (11)
  • [5] SparseLIF: High-Performance Sparse LiDAR-Camera Fusion for 3D Object Detection
    Zhang, Hongcheng
    Liang, Liu
    Zeng, Pengxin
    Song, Xiao
    Wang, Zhe
    COMPUTER VISION-ECCV 2024, PT XXXV, 2025, 15093 : 109 - 128
  • [6] FS-Net: LiDAR-Camera Fusion With Matched Scale for 3D Object Detection in Autonomous Driving
    Zhang, Lei
    Li, Xu
    Tang, Kaichen
    Jiang, Yunzhe
    Yang, Liu
    Zhang, Yonggang
    Chen, Xianyi
    IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS, 2023, 24 (11) : 12154 - 12165
  • [7] CramNet: Camera-Radar Fusion with Ray-Constrained Cross-Attention for Robust 3D Object Detection
    Hwang, Jyh-Jing
    Kretzschmar, Henrik
    Manela, Joshua
    Rafferty, Sean
    Armstrong-Crews, Nicholas
    Chen, Tiffany
    Anguelov, Dragomir
    COMPUTER VISION, ECCV 2022, PT XXXVIII, 2022, 13698 : 388 - 405
  • [8] Multi-Scale Spatial Transformer Network for LiDAR-Camera 3D Object Detection
    Wang, Zhifan
    Zhang, Xiaohong
    Wang, Shidong
    Xin, Tong
    Zhang, Haofeng
    Lu, Jianfeng
    2021 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS (IJCNN), 2021,
  • [9] BEV-CFKT: A LiDAR-camera cross-modality-interaction fusion and knowledge transfer framework with transformer for BEV 3D object detection
    Wei, Ming
    Li, Jiachen
    Kang, Hongyi
    Huang, Yijie
    Lu, Jun-Guo
    NEUROCOMPUTING, 2024, 582
  • [10] Enhanced Object Detection in Autonomous Vehicles through LiDAR-Camera Sensor Fusion
    Dai, Zhongmou
    Guan, Zhiwei
    Chen, Qiang
    Xu, Yi
    Sun, Fengyi
    WORLD ELECTRIC VEHICLE JOURNAL, 2024, 15 (07):