Self-Supervised 3D Point Cloud Detection with Depthwise Separable Masked Autoencoders
Three-dimensional object detection from LiDAR point clouds is essential for autonomous driving, yet existing methods typically rely on costly and extensively annotated datasets. Self-supervised masked autoencoders (MAE) provide a promising alternative, but effectively capturing local geometric structures while maintaining computational efficiency remains challenging. This paper presents DS-MAE, a depthwise separable masked autoencoder for self-supervised 3D point cloud detection. DS-MAE introduces a pyramidal transformer encoder with depthwise separable attention to enhance local geometric feature extraction while reducing computational cost. A depthwise separable generative decoder is further designed for multi-scale masked feature reconstruction, while a density-aware reconstruction loss accounts for the non-uniform point density of LiDAR observations. Experiments on the KITTI, Waymo Open Dataset (WOD), and a small-scale subset of the ONCE dataset demonstrate the effectiveness of DS-MAE, achieving 67.95%, 67.70%, and 58.92% mAP, respectively. On WOD, the reported performance is achieved using only 20% of the labeled training data, demonstrating the effectiveness of DS-MAE under limited supervision.
Authors
- Sirui Guo (ORCID: https://orcid.org/0000-0002-4708-5959)
- Jie Cao (ORCID: https://orcid.org/0000-0001-8376-7669)
- Yaqian Ning
- Zhaoyang Li (ORCID: https://orcid.org/0000-0002-3388-0388)
- Jiayi Zhou (ORCID: https://orcid.org/0000-0001-9265-5933)
- Pengcheng Ji
Institutions
- Changchun University of Science and Technology (CN)
- Beijing Institute of Technology (CN)
Publication Details
- Journal
- Remote Sensing
- Published
- 2026-09-13
- DOI
- https://doi.org/10.3390/rs18183148
- Primary Topic
- 3D Shape Modeling and Analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00