BAFusion: Bidirectional Attention Fusion for 3D Object Detection Based on LiDAR and Camera

被引:4
|
作者
Liu, Min [1 ]
Jia, Yuanjun [2 ]
Lyu, Youhao [1 ]
Dong, Qi [2 ]
Yang, Yanyu [2 ]
机构
[1] Univ Sci & Technol China, Inst Adv Technol, Hefei 230088, Peoples R China
[2] China Acad Elect & Informat Technol, Beijing 100041, Peoples R China
关键词
3D object detection; LiDAR-camera fusion; cross attention;
D O I
10.3390/s24144718
中图分类号
O65 [分析化学];
学科分类号
070302 ; 081704 ;
摘要
3D object detection is a challenging and promising task for autonomous driving and robotics, benefiting significantly from multi-sensor fusion, such as LiDAR and cameras. Conventional methods for sensor fusion rely on a projection matrix to align the features from LiDAR and cameras. However, these methods often suffer from inadequate flexibility and robustness, leading to lower alignment accuracy under complex environmental conditions. Addressing these challenges, in this paper, we propose a novel Bidirectional Attention Fusion module, named BAFusion, which effectively fuses the information from LiDAR and cameras using cross-attention. Unlike the conventional methods, our BAFusion module can adaptively learn the cross-modal attention weights, making the approach more flexible and robust. Moreover, drawing inspiration from advanced attention optimization techniques in 2D vision, we developed the Cross Focused Linear Attention Fusion Layer (CFLAF Layer) and integrated it into our BAFusion pipeline. This layer optimizes the computational complexity of attention mechanisms and facilitates advanced interactions between image and point cloud data, showcasing a novel approach to addressing the challenges of cross-modal attention calculations. We evaluated our method on the KITTI dataset using various baseline networks, such as PointPillars, SECOND, and Part-A2, and demonstrated consistent improvements in 3D object detection performance over these baselines, especially for smaller objects like cyclists and pedestrians. Our approach achieves competitive results on the KITTI benchmark.
引用
收藏
页数:26
相关论文
共 33 条
  • [21] ICAFusion: Iterative cross-attention guided feature fusion for multispectral object detection
    Shen, Jifeng
    Chen, Yifei
    Liu, Yue
    Zuo, Xin
    Fan, Heng
    Yang, Wankou
    PATTERN RECOGNITION, 2024, 145
  • [22] Background-Aware Cross-Attention Multiscale Fusion for Multispectral Object Detection
    Guo, Runze
    Guo, Xiaojun
    Sun, Xiaoyong
    Zhou, Peida
    Sun, Bei
    Su, Shaojing
    REMOTE SENSING, 2024, 16 (21)
  • [23] CFRNet: Cross-Attention-Based Fusion and Refinement Network for Enhanced RGB-T Salient Object Detection
    Deng, Biao
    Liu, Di
    Cao, Yang
    Liu, Hong
    Yan, Zhiguo
    Chen, Hu
    SENSORS, 2024, 24 (22)
  • [24] SCA-PVNet: Self-and-cross attention based aggregation of point cloud and multi-view for 3D object retrieval
    Lin, Dongyun
    Cheng, Yi
    Guo, Aiyuan
    Mao, Shangbo
    Li, Yiqun
    KNOWLEDGE-BASED SYSTEMS, 2024, 296
  • [25] DACFusion: Dual Asymmetric Cross-Attention guided feature fusion for multispectral object detection
    Qian, Jingchen
    Qiao, Baiyou
    Zhang, Yuekai
    Liu, Tongyan
    Wang, Shuo
    Wu, Gang
    Han, Donghong
    NEUROCOMPUTING, 2025, 635
  • [26] Towards Camera-LIDAR Fusion-Based Terrain Modelling for Planetary Surfaces: Review and Analysis
    Shaukat, Affan
    Blacker, Peter C.
    Spiteri, Conrad
    Gao, Yang
    SENSORS, 2016, 16 (11):
  • [27] 3D lymphoma segmentation on PET/CT images via multi-scale information fusion with cross-attention
    Huang, Huan
    Qiu, Liheng
    Yang, Shenmiao
    Li, Longxi
    Nan, Jiaofen
    Li, Yanting
    Han, Chuang
    Zhu, Fubao
    Zhao, Chen
    Zhou, Weihua
    MEDICAL PHYSICS, 2025,
  • [28] Cross Attention-Based Multi-Scale Convolutional Fusion Network for Hyperspectral and LiDAR Joint Classification
    Ge, Haimiao
    Wang, Liguo
    Pan, Haizhu
    Liu, Yanzhong
    Li, Cheng
    Lv, Dan
    Ma, Huiyu
    REMOTE SENSING, 2024, 16 (21)
  • [29] VioNets: efficient multi-modal fusion method based on bidirectional gate recurrent unit and cross-attention graph convolutional network for video violence detection
    Liang, Wuyan
    Xu, Xiaolong
    Fu, Xiao
    JOURNAL OF ELECTRONIC IMAGING, 2023, 32 (02)
  • [30] 3D Rock and Pothole Detection in Desert for the Wild Navigation
    Yuan, Yibo
    Xue, Jianru
    Hu, Tao
    Zheng, Bo
    Zhou, Yang
    Song, Wenjie
    Cao, Tao
    INTELLIGENT ROBOTICS AND APPLICATIONS, ICIRA 2024, PT IX, 2025, 15209 : 441 - 455