HSCM-Lane: A ResNet-based encoder-decoder architecture with multi-level shifted-window context modeling for pixel-wise lane segmentation

Pixel-wise lane segmentation provides dense geometric cues for lane keeping, road-structure inference, and local trajectory planning in advanced driver-assistance systems and autonomous vehicles. The objective is to improve binary lane-mask segmentation under sparse, thin, low-contrast, occluded, and discontinuous lane-marking conditions while retaining a compact real-time configuration. To this end, we propose HSCM-Lane, a ResNet-based encoder-decoder architecture enhanced by a Hierarchical Swin Context Module (HSCM). HSCM organizes shifted-window context modeling at the 1/4, 1/8, and 1/16 encoder feature levels, so that context-enhanced features are propagated through subsequent residual stages and reused by a lightweight decoder with projected static-sum fusion. For backbone-capacity analysis, the architecture is evaluated as HSCM-Lane-S (Small, ResNet18), HSCM-Lane-M (Medium, ResNet34), and HSCM-Lane-L (Large, ResNet50). These configurations differ only in the ResNet backbone, while the HSCM placement, window size, depth, projected static-sum fusion strategy, and Focal-Tversky objective are kept fixed. On the BDD100K validation set, HSCM-Lane-S achieves 34.44% lane-class intersection over union (IoU), 65.00% lane recall, and 82.12% balanced accuracy, compared with 33.00%, 62.48%, and 80.38% for the no-HSCM counterpart. Relative to the no-HSCM counterpart, these values represent gains of 1.44 percentage points in lane IoU, 2.52 points in lane recall, and 1.74 points in balanced accuracy. The lane-IoU gain corresponds to a 4.36% relative improvement. Scaling the backbone yields 34.86% lane IoU for HSCM-Lane-M and 34.98% lane IoU for HSCM-Lane-L, showing a gradual quality-cost trade-off. Under a derived pixel-level out-of-domain evaluation protocol on TuSimple, HSCM-Lane-S, HSCM-Lane-M, and HSCM-Lane-L achieve 27.2%, 27.6%, and 27.6% lane IoU, respectively, without target-domain fine-tuning. These results indicate that multi-level shifted-window context modeling improves the no-HSCM counterpart and yields competitive lane-mask segmentation results under these protocols. Future work will address temporal, instance-level, multimodal, planning, and safety validation.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-10-07
DOI
https://doi.org/10.1371/journal.pone.0359959
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

HSCM-Lane: A ResNet-based encoder-decoder architecture with multi-level shifted-window context modeling for pixel-wise lane segmentation

Tuong Le, Thanh Nguyên Vũ, Thanh Hien Vu, Quang Toai Ton
PLoS ONE
Advanced Neural Network Applications
article

HSCM-Lane: A ResNet-based encoder-decoder architecture with multi-level shifted-window context modeling for pixel-wise lane segmentation

Tuong Le, Thanh Nguyên Vũ, Thanh Hien Vu, Quang Toai Ton
article en

Abstract

Pixel-wise lane segmentation provides dense geometric cues for lane keeping, road-structure inference, and local trajectory planning in advanced driver-assistance systems and autonomous vehicles. The objective is to improve binary lane-mask segmentation under sparse, thin, low-contrast, occluded, and discontinuous lane-marking conditions while retaining a compact real-time configuration. To this end, we propose HSCM-Lane, a ResNet-based encoder-decoder architecture enhanced by a Hierarchical Swin Context Module (HSCM). HSCM organizes shifted-window context modeling at the 1/4, 1/8, and 1/16 encoder feature levels, so that context-enhanced features are propagated through subsequent residual stages and reused by a lightweight decoder with projected static-sum fusion. For backbone-capacity analysis, the architecture is evaluated as HSCM-Lane-S (Small, ResNet18), HSCM-Lane-M (Medium, ResNet34), and HSCM-Lane-L (Large, ResNet50). These configurations differ only in the ResNet backbone, while the HSCM placement, window size, depth, projected static-sum fusion strategy, and Focal-Tversky objective are kept fixed. On the BDD100K validation set, HSCM-Lane-S achieves 34.44% lane-class intersection over union (IoU), 65.00% lane recall, and 82.12% balanced accuracy, compared with 33.00%, 62.48%, and 80.38% for the no-HSCM counterpart. Relative to the no-HSCM counterpart, these values represent gains of 1.44 percentage points in lane IoU, 2.52 points in lane recall, and 1.74 points in balanced accuracy. The lane-IoU gain corresponds to a 4.36% relative improvement. Scaling the backbone yields 34.86% lane IoU for HSCM-Lane-M and 34.98% lane IoU for HSCM-Lane-L, showing a gradual quality-cost trade-off. Under a derived pixel-level out-of-domain evaluation protocol on TuSimple, HSCM-Lane-S, HSCM-Lane-M, and HSCM-Lane-L achieve 27.2%, 27.6%, and 27.6% lane IoU, respectively, without target-domain fine-tuning. These results indicate that multi-level shifted-window context modeling improves the no-HSCM counterpart and yields competitive lane-mask segmentation results under these protocols. Future work will address temporal, instance-level, multimodal, planning, and safety validation.

PLoS ONEVol. 21(10)
Ho Chi Minh City University of Foreign Languages and Information Technology (VN), Ho Chi Minh City University of Industry and Trade (VN), HUTECH University
Openalex Percentile: Top 15%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.