Rethinking Masked Representation Learning for 3D Point Cloud Understanding

被引：0

作者：

Wang, Chuxin ^{[1
,2
]}

Zha, Yixin ^{[1
,2
]}

He, Jianfeng ^{[1
,2
]}

Yang, Wenfei ^{[1
,2
]}

Zhang, Tianzhu ^{[1
,2
]}

机构：

[1] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230027, Peoples R China

[2] Univ Sci & Technol China, Deep Space Explorat Lab, Hefei 230027, Peoples R China

来源：

IEEE TRANSACTIONS ON IMAGE PROCESSING | 2025年 / 34卷

关键词：

Point cloud compression; Semantics; Feature extraction; Three-dimensional displays; Representation learning; Solid modeling; Prototypes; Shape; Nearest neighbor methods; Image reconstruction; Self-supervised point cloud representation learning; optimal transport; and part modeling; NETWORK;

D O I：

10.1109/TIP.2024.3520008

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Self-supervised point cloud representation learning aims to acquire robust and general feature representations from unlabeled data. Recently, masked point modeling-based methods have shown significant performance improvements for point cloud understanding, yet these methods rely on overlapping grouping strategies (k-nearest neighbor algorithm) resulting in early leakage of structural information of mask groups, and overlook the semantic modeling of object components resulting in parts with the same semantics having obvious feature differences due to position differences. In this work, we rethink grouping strategies and pretext tasks that are more suitable for self-supervised point cloud representation learning and propose a novel hierarchical masked representation learning method, including an optimal transport-based hierarchical grouping strategy, a prototype-based part modeling module, and a hierarchical attention encoder. The proposed method enjoys several merits. First, the proposed grouping strategy partitions the point cloud into non-overlapping groups, eliminating the early leakage of structural information in the masked groups. Second, the proposed prototype-based part modeling module dynamically models different object components, ensuring feature consistency on parts with the same semantics. Extensive experiments on four downstream tasks demonstrate that our method surpasses state-of-the-art 3D representation learning methods. Furthermore, Comprehensive ablation studies and visualizations demonstrate the effectiveness of the proposed modules.

引用

页码：247 / 262

页数：16

共 50 条

[1] Masked Autoencoders in 3D Point Cloud Representation Learning
Jiang, Jincen
Lu, Xuequan
Zhao, Lizhi
Dazeley, Richard
Wang, Meili
IEEE TRANSACTIONS ON MULTIMEDIA, 2025, 27 : 820 - 831
[2] Masked Structural Point Cloud Modeling to Learning 3D Representation
Yamada, Ryosuke
Tadokoro, Ryu
Qiu, Yue
Kataoka, Hirokatsu
Satoh, Yutaka
IEEE ACCESS, 2024, 12 : 142291 - 142305
[3] PatchMixing Masked Autoencoders for 3D Point Cloud Self-Supervised Learning
Lin, Chengxing
Xu, Wenju
Zhu, Jian
Nie, Yongwei
Cai, Ruichu
Xu, Xuemiao
IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2024, 34 (10) : 9882 - 9897
[4] Learning 3D Shape Latent for Point Cloud Completion
Chen, Zhikai
Long, Fuchen
Qiu, Zhaofan
Yao, Ting
Zhou, Wengang
Luo, Jiebo
Mei, Tao
IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 : 8717 - 8729
[5] Flattening-Net: Deep Regular 2D Representation for 3D Point Cloud Analysis
Zhang, Qijian
Hou, Junhui
Qian, Yue
Zeng, Yiming
Zhang, Juyong
He, Ying
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2023, 45 (08) : 9726 - 9742
[6] Joint representation learning for text and 3D point cloud
Huang, Rui
Pan, Xuran
Zheng, Henry
Jiang, Haojun
Xie, Zhifeng
Wu, Cheng
Song, Shiji
Huang, Gao
PATTERN RECOGNITION, 2024, 147
[7] Geometric Invariant Representation Learning for 3D Point Cloud
Li, Zongmin
Zhang, Yupeng
Bai, Yun
2021 IEEE 33RD INTERNATIONAL CONFERENCE ON TOOLS WITH ARTIFICIAL INTELLIGENCE (ICTAI 2021), 2021, : 1480 - 1485
[8] WSSIC-Net: Weakly-Supervised Semantic Instance Completion of 3D Point Cloud Scenes
Fu, Zhiheng
Guo, Yulan
Chen, Minglin
Hu, Qingyong
Laga, Hamid
Boussaid, Farid
Bennamoun, Mohammed
IEEE TRANSACTIONS ON IMAGE PROCESSING, 2025, 34 : 2008 - 2019
[9] Exploring Self-Supervised Learning for 3D Point Cloud Registration
Yuan, Mingzhi
Huang, Qiao
Shen, Ao
Huang, Xiaoshui
Wang, Manning
IEEE ROBOTICS AND AUTOMATION LETTERS, 2025, 10 (01): : 25 - 31
[10] LinK3D: Linear Keypoints Representation for 3D LiDAR Point Cloud
Cui, Yunge
Zhang, Yinlong
Dong, Jiahua
Sun, Haibo
Chen, Xieyuanli
Zhu, Feng
IEEE ROBOTICS AND AUTOMATION LETTERS, 2024, 9 (03) : 2128 - 2135

← 1 2 3 4 5 →