Dual stream knowledge distillation for 3D LiDAR segmentation model with large sparse multi-directional convolution

Semantic segmentation of LiDAR point clouds is a fundamental task in 3D scene understanding, with broad applications in autonomous driving, robotics, and urban mapping. This paper presents a novel framework that couples 3D large sparse multi-directional convolution with a training-time dual-stream knowledge distillation strategy. The backbone employs dynamic large sparse kernels and multi-directional convolution to create large effective receptive fields, while radially non-uniform cylindrical voxelization balances point density across the scene. During training, knowledge is distilled from two auxiliary modalities—2D RGB images and 1D space-filling-curve serialized point clouds—into the 3D backbone via a multi-layer bidirectional polarity-aware linear cross-attention mechanism, adding zero inference cost. Experiments on SemanticKITTI and nuScenes achieve mean IoU of 74.4% and 83.5%, respectively.

Authors

Institutions

Publication Details

Journal
Complex & Intelligent Systems
Published
2026-09-28
DOI
https://doi.org/10.1007/s40747-026-02523-w
Primary Topic
Sparse and Compressive Sensing Techniques
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Dual stream knowledge distillation for 3D LiDAR segmentation model with large sparse multi-directional convolution

Ye Gu, Huancheng Xiao, Gengliang Chen, Lixin Liang
Complex & Intelligent Systems
Sparse and Compressive Sensing Techniques
article

Dual stream knowledge distillation for 3D LiDAR segmentation model with large sparse multi-directional convolution

Ye Gu, Huancheng Xiao, Gengliang Chen, Lixin Liang
article en

Abstract

Semantic segmentation of LiDAR point clouds is a fundamental task in 3D scene understanding, with broad applications in autonomous driving, robotics, and urban mapping. This paper presents a novel framework that couples 3D large sparse multi-directional convolution with a training-time dual-stream knowledge distillation strategy. The backbone employs dynamic large sparse kernels and multi-directional convolution to create large effective receptive fields, while radially non-uniform cylindrical voxelization balances point density across the scene. During training, knowledge is distilled from two auxiliary modalities—2D RGB images and 1D space-filling-curve serialized point clouds—into the 3D backbone via a multi-layer bidirectional polarity-aware linear cross-attention mechanism, adding zero inference cost. Experiments on SemanticKITTI and nuScenes achieve mean IoU of 74.4% and 83.5%, respectively.

Complex & Intelligent Systems
Shenzhen Technology University (CN)
Shenzhen Science and Technology Innovation Program
Openalex Percentile: Top 14%
Sparse and Compressive Sensing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Dual stream knowledge distillation for 3D LiDAR segmentation model with large sparse multi-directional convolution — Ye Gu, Huancheng Xiao, et al. · Complex & Intelligent Systems (2026) | TGRS Research Map | TGRS