BAFusion: Bidirectional Attention Fusion for 3D Object Detection Based on LiDAR and Camera

被引:4
|
作者
Liu, Min [1 ]
Jia, Yuanjun [2 ]
Lyu, Youhao [1 ]
Dong, Qi [2 ]
Yang, Yanyu [2 ]
机构
[1] Univ Sci & Technol China, Inst Adv Technol, Hefei 230088, Peoples R China
[2] China Acad Elect & Informat Technol, Beijing 100041, Peoples R China
关键词
3D object detection; LiDAR-camera fusion; cross attention;
D O I
10.3390/s24144718
中图分类号
O65 [分析化学];
学科分类号
070302 ; 081704 ;
摘要
3D object detection is a challenging and promising task for autonomous driving and robotics, benefiting significantly from multi-sensor fusion, such as LiDAR and cameras. Conventional methods for sensor fusion rely on a projection matrix to align the features from LiDAR and cameras. However, these methods often suffer from inadequate flexibility and robustness, leading to lower alignment accuracy under complex environmental conditions. Addressing these challenges, in this paper, we propose a novel Bidirectional Attention Fusion module, named BAFusion, which effectively fuses the information from LiDAR and cameras using cross-attention. Unlike the conventional methods, our BAFusion module can adaptively learn the cross-modal attention weights, making the approach more flexible and robust. Moreover, drawing inspiration from advanced attention optimization techniques in 2D vision, we developed the Cross Focused Linear Attention Fusion Layer (CFLAF Layer) and integrated it into our BAFusion pipeline. This layer optimizes the computational complexity of attention mechanisms and facilitates advanced interactions between image and point cloud data, showcasing a novel approach to addressing the challenges of cross-modal attention calculations. We evaluated our method on the KITTI dataset using various baseline networks, such as PointPillars, SECOND, and Part-A2, and demonstrated consistent improvements in 3D object detection performance over these baselines, especially for smaller objects like cyclists and pedestrians. Our approach achieves competitive results on the KITTI benchmark.
引用
收藏
页数:26
相关论文
共 50 条
  • [21] FGFusion: Fine-Grained Lidar-Camera Fusion for 3D Object Detection
    Yin, Zixuan
    Sun, Han
    Liu, Ningzhong
    Zhou, Huiyu
    Shen, Jiaquan
    PATTERN RECOGNITION AND COMPUTER VISION, PRCV 2023, PT III, 2024, 14427 : 505 - 517
  • [22] Deep structural information fusion for 3D object detection on LiDAR-camera system
    An, Pei
    Liang, Junxiong
    Yu, Kun
    Fang, Bin
    Ma, Jie
    COMPUTER VISION AND IMAGE UNDERSTANDING, 2022, 214
  • [23] 3D Object Detection with LiDAR Based on Multi-Attention Mechanism
    Cao, Jie
    Peng, Yiqiang
    Fan, Likang
    Mo, Lingfan
    Wang, Longfei
    LASER & OPTOELECTRONICS PROGRESS, 2025, 62 (04)
  • [24] Fast-CLOCs: Fast Camera-LiDAR Object Candidates Fusion for 3D Object Detection
    Pang, Su
    Morris, Daniel
    Radha, Hayder
    2022 IEEE WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV 2022), 2022, : 3747 - 3756
  • [25] 3D object detection based on image and LIDAR fusion for autonomous driving
    Chen G.
    Yi H.
    Mao Z.
    International Journal of Vehicle Information and Communication Systems, 2023, 8 (03) : 237 - 251
  • [26] Research and application of indoor 3D object detection based on Lidar and Monocular camera
    Jiang, Yi
    Sun, Bingyu
    2024 5TH INTERNATIONAL CONFERENCE ON COMPUTER ENGINEERING AND APPLICATION, ICCEA 2024, 2024, : 1119 - 1123
  • [27] CaLiJD: Camera and LiDAR Joint Contender for 3D Object Detection
    Lyu, Jiahang
    Qi, Yongze
    You, Suilian
    Meng, Jin
    Meng, Xin
    Kodagoda, Sarath
    Wang, Shifeng
    REMOTE SENSING, 2024, 16 (23)
  • [28] PLC-Fusion: Perspective-Based Hierarchical and Deep LiDAR Camera Fusion for 3D Object Detection in Autonomous Vehicles
    Mushtaq, Husnain
    Deng, Xiaoheng
    Azhar, Fizza
    Ali, Mubashir
    Sherazi, Hafiz Husnain Raza
    INFORMATION, 2024, 15 (11)
  • [29] DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection
    Li, Yingwei
    Yu, Adams Wei
    Meng, Tianjian
    Caine, Ben
    Ngiam, Jiquan
    Peng, Daiyi
    Shen, Junyang
    Lu, Yifeng
    Zhou, Denny
    Le, Quoc, V
    Yuille, Alan
    Tan, Mingxing
    2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, : 17161 - 17170
  • [30] SparseLIF: High-Performance Sparse LiDAR-Camera Fusion for 3D Object Detection
    Zhang, Hongcheng
    Liang, Liu
    Zeng, Pengxin
    Song, Xiao
    Wang, Zhe
    COMPUTER VISION-ECCV 2024, PT XXXV, 2025, 15093 : 109 - 128