Dual stream knowledge distillation for 3D LiDAR segmentation model with large sparse multi-directional convolution
Semantic segmentation of LiDAR point clouds is a fundamental task in 3D scene understanding, with broad applications in autonomous driving, robotics, and urban mapping. This paper presents a novel framework that couples 3D large sparse multi-directional convolution with a training-time dual-stream knowledge distillation strategy. The backbone employs dynamic large sparse kernels and multi-directional convolution to create large effective receptive fields, while radially non-uniform cylindrical voxelization balances point density across the scene. During training, knowledge is distilled from two auxiliary modalities—2D RGB images and 1D space-filling-curve serialized point clouds—into the 3D backbone via a multi-layer bidirectional polarity-aware linear cross-attention mechanism, adding zero inference cost. Experiments on SemanticKITTI and nuScenes achieve mean IoU of 74.4% and 83.5%, respectively.
Authors
- Ye Gu (ORCID: https://orcid.org/0000-0002-4008-2042)
- Huancheng Xiao
- Gengliang Chen
- Lixin Liang
Institutions
- Shenzhen Technology University (CN)
Publication Details
- Journal
- Complex & Intelligent Systems
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1007/s40747-026-02523-w
- Primary Topic
- Sparse and Compressive Sensing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Shenzhen Science and Technology Innovation Program