SparseDet: A Simple and Effective Framework for Fully Sparse LiDAR-Based 3-D Object Detection

被引：1

作者：

Liu, Lin ^{[1
]}

Song, Ziying ^{[1
]}

Xia, Qiming ^{[2
]}

Jia, Feiyang ^{[1
]}

Jia, Caiyan ^{[1
]}

Yang, Lei ^{[3
,4
]}

Gong, Yan ^{[5
]}

Pan, Hongyu ^{[6
]}

机构：

[1] Beijing Jiaotong Univ, Sch Comp Sci & Technol, Beijing Key Lab Traff Data Anal & Min, Beijing 100044, Peoples R China

[2] Xiamen Univ, Fujian Key Lab Sensing & Comp Smart Cities, Xiamen 361005, Fujian, Peoples R China

[3] Tsinghua Univ, State Key Lab Intelligent Green Vehicle & Mobil, Beijing 100084, Peoples R China

[4] Tsinghua Univ, Sch Vehicle & Mobil, Beijing 100084, Peoples R China

[5] JD Logist, Autonomous Driving Dept X Div, Beijing 101111, Peoples R China

[6] Horizon Robot, Beijing 100190, Peoples R China

来源：

IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING | 2024年 / 62卷

关键词：

Feature extraction; Three-dimensional displays; Point cloud compression; Detectors; Aggregates; Object detection; Computational efficiency; 3-D object detection; feature aggregation; sparse detectors;

D O I：

10.1109/TGRS.2024.3468394

中图分类号：

P3 [地球物理学]; P59 [地球化学];

学科分类号：

0708 ; 070902 ;

摘要：

LiDAR-based sparse 3-D object detection plays a crucial role in autonomous driving applications due to its computational efficiency advantages. Existing methods either use the features of a single central voxel as an object proxy or treat an aggregated cluster of foreground points as an object proxy. However, the former cannot aggregate contextual information, resulting in insufficient information expression in object proxies. The latter relies on multistage pipelines and auxiliary tasks, which reduce the inference speed. To maintain the efficiency of the sparse framework while fully aggregating contextual information, in this work, we propose SparseDet that designs sparse queries as object proxies. It introduces two key modules: the local multiscale feature aggregation (LMFA) module and the global feature aggregation (GFA) module, aiming to fully capture the contextual information, thereby enhancing the ability of the proxies to represent objects. The LMFA module achieves feature fusion across different scales for sparse key voxels via coordinate transformations and using nearest neighbor relationships to capture object-level details and local contextual information, whereas the GFA module uses self-attention mechanisms to selectively aggregate the features of the key voxels across the entire scene for capturing scene-level contextual information. Experiments on nuScenes and KITTI demonstrate the effectiveness of our method. Specifically, SparseDet surpasses the previous best sparse detector VoxelNeXt (a typical method using voxels as object proxies) by 2.2% mean average precision (mAP) with 13.5 frames/s on nuScenes and outperforms VoxelNeXt by 1.12% AP(3-D) on hard level tasks with 17.9 frames/s on KITTI. What is more, not only the mAP of SparseDet exceeds that of FSDV2 (a classical method using clusters of foreground points as object proxies) but also its inference speed is 1.3 times faster than FSDV2 on the nuScenes test set. The code has been released in https://github.com/liulin813/SparseDet.git.

引用

页数：14

共 50 条

[31] LiDAR-Based Multisensor Fusion With 3-D Digital Maps for High-Precision Positioning
Mounier, Eslam
Elhabiby, Mohamed
Korenberg, Michael
Noureldin, Aboelmagd
IEEE INTERNET OF THINGS JOURNAL, 2025, 12 (06): : 7209 - 7224
[32] Registration for 3-D LiDAR Datasets Using Pyramid Reference Object
Song, Wei
Li, Dechao
Sun, Su
Xu, Xinghui
Zu, Guidong
IEEE TRANSACTIONS ON INSTRUMENTATION AND MEASUREMENT, 2023, 72
[33] LiDAR-Based Optimized Normal Distribution Transform Localization on 3-D Map for Autonomous Navigation
Thakur, Abhishek
Rajalakshmi, P.
IEEE OPEN JOURNAL OF INSTRUMENTATION AND MEASUREMENT, 2024, 3
[34] Noise Model-Based Line Segmentation for Plane Extraction in Sparse 3-D LiDAR Data
He, Linkun
Li, Bofeng
Chen, Guang'e
IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 2024, 62 : 1 - 15
[35] Efficient 3D Object Detection Based on Pseudo-LiDAR Representation
Meng, Haitao
Li, Changcai
Chen, Gang
Chen, Long
Knoll, Alois
IEEE TRANSACTIONS ON INTELLIGENT VEHICLES, 2024, 9 (01): : 1953 - 1964
[36] A Novel Sparse Geometric 3-D LiDAR Odometry Approach
Liang, Shuang
Cao, Zhiqiang
Guan, Peiyu
Wang, Chengpeng
Yu, Junzhi
Wang, Shuo
IEEE SYSTEMS JOURNAL, 2021, 15 (01): : 1390 - 1400
[37] Fast LiDAR R-CNN: Residual Relation-Aware Region Proposal Networks for Multiclass 3-D Object Detection
Wen, Lihua
Jo, Kang-Hyun
IEEE SENSORS JOURNAL, 2022, 22 (12) : 12323 - 12331
[38] ODD-M3D: Object-Wise Dense Depth Estimation for Monocular 3-D Object Detection
Park, Chanyeong
Kim, Heegwang
Jang, Junbo
Paik, Joonki
IEEE TRANSACTIONS ON CONSUMER ELECTRONICS, 2024, 70 (01) : 646 - 655
[39] SP-Det: Leveraging Saliency Prediction for Voxel-Based 3D Object Detection in Sparse Point Cloud
An, Pei
Duan, Yucong
Huang, Yuliang
Ma, Jie
Chen, Yanfei
Wang, Liheng
Yang, You
Liu, Qiong
IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 (2795-2808) : 2795 - 2808
[40] 3D-DFM: Anchor-Free Multimodal 3-D Object Detection With Dynamic Fusion Module for Autonomous Driving
Lin, Chunmian
Tian, Daxin
Duan, Xuting
Zhou, Jianshan
Zhao, Dezong
Cao, Dongpu
IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2023, 34 (12) : 10812 - 10822

← 1 2 3 4 5 →