Self-Supervised 3D Point Cloud Detection with Depthwise Separable Masked Autoencoders

Three-dimensional object detection from LiDAR point clouds is essential for autonomous driving, yet existing methods typically rely on costly and extensively annotated datasets. Self-supervised masked autoencoders (MAE) provide a promising alternative, but effectively capturing local geometric structures while maintaining computational efficiency remains challenging. This paper presents DS-MAE, a depthwise separable masked autoencoder for self-supervised 3D point cloud detection. DS-MAE introduces a pyramidal transformer encoder with depthwise separable attention to enhance local geometric feature extraction while reducing computational cost. A depthwise separable generative decoder is further designed for multi-scale masked feature reconstruction, while a density-aware reconstruction loss accounts for the non-uniform point density of LiDAR observations. Experiments on the KITTI, Waymo Open Dataset (WOD), and a small-scale subset of the ONCE dataset demonstrate the effectiveness of DS-MAE, achieving 67.95%, 67.70%, and 58.92% mAP, respectively. On WOD, the reported performance is achieved using only 20% of the labeled training data, demonstrating the effectiveness of DS-MAE under limited supervision.

Authors

Institutions

Publication Details

Journal
Remote Sensing
Published
2026-09-13
DOI
https://doi.org/10.3390/rs18183148
Primary Topic
3D Shape Modeling and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Self-Supervised 3D Point Cloud Detection with Depthwise Separable Masked Autoencoders

Sirui Guo, Jie Cao, Yaqian Ning, Zhaoyang Li et al.
Remote Sensing
3D Shape Modeling and Analysis
article

Self-Supervised 3D Point Cloud Detection with Depthwise Separable Masked Autoencoders

Sirui Guo, Jie Cao, Yaqian Ning, Zhaoyang Li, Jiayi Zhou, Pengcheng Ji
article en

Abstract

Three-dimensional object detection from LiDAR point clouds is essential for autonomous driving, yet existing methods typically rely on costly and extensively annotated datasets. Self-supervised masked autoencoders (MAE) provide a promising alternative, but effectively capturing local geometric structures while maintaining computational efficiency remains challenging. This paper presents DS-MAE, a depthwise separable masked autoencoder for self-supervised 3D point cloud detection. DS-MAE introduces a pyramidal transformer encoder with depthwise separable attention to enhance local geometric feature extraction while reducing computational cost. A depthwise separable generative decoder is further designed for multi-scale masked feature reconstruction, while a density-aware reconstruction loss accounts for the non-uniform point density of LiDAR observations. Experiments on the KITTI, Waymo Open Dataset (WOD), and a small-scale subset of the ONCE dataset demonstrate the effectiveness of DS-MAE, achieving 67.95%, 67.70%, and 58.92% mAP, respectively. On WOD, the reported performance is achieved using only 20% of the labeled training data, demonstrating the effectiveness of DS-MAE under limited supervision.

Remote SensingVol. 18(18)
Changchun University of Science and Technology (CN), Beijing Institute of Technology (CN)
Industry, innovation and infrastructure
Openalex Percentile: Top 13%
3D Shape Modeling and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Self-Supervised 3D Point Cloud Detection with Depthwise Separable Masked Autoencoders — Sirui Guo, Jie Cao, et al. · Remote Sensing (2026) | TGRS Research Map | TGRS