HSCM-Lane: A ResNet-based encoder-decoder architecture with multi-level shifted-window context modeling for pixel-wise lane segmentation
Pixel-wise lane segmentation provides dense geometric cues for lane keeping, road-structure inference, and local trajectory planning in advanced driver-assistance systems and autonomous vehicles. The objective is to improve binary lane-mask segmentation under sparse, thin, low-contrast, occluded, and discontinuous lane-marking conditions while retaining a compact real-time configuration. To this end, we propose HSCM-Lane, a ResNet-based encoder-decoder architecture enhanced by a Hierarchical Swin Context Module (HSCM). HSCM organizes shifted-window context modeling at the 1/4, 1/8, and 1/16 encoder feature levels, so that context-enhanced features are propagated through subsequent residual stages and reused by a lightweight decoder with projected static-sum fusion. For backbone-capacity analysis, the architecture is evaluated as HSCM-Lane-S (Small, ResNet18), HSCM-Lane-M (Medium, ResNet34), and HSCM-Lane-L (Large, ResNet50). These configurations differ only in the ResNet backbone, while the HSCM placement, window size, depth, projected static-sum fusion strategy, and Focal-Tversky objective are kept fixed. On the BDD100K validation set, HSCM-Lane-S achieves 34.44% lane-class intersection over union (IoU), 65.00% lane recall, and 82.12% balanced accuracy, compared with 33.00%, 62.48%, and 80.38% for the no-HSCM counterpart. Relative to the no-HSCM counterpart, these values represent gains of 1.44 percentage points in lane IoU, 2.52 points in lane recall, and 1.74 points in balanced accuracy. The lane-IoU gain corresponds to a 4.36% relative improvement. Scaling the backbone yields 34.86% lane IoU for HSCM-Lane-M and 34.98% lane IoU for HSCM-Lane-L, showing a gradual quality-cost trade-off. Under a derived pixel-level out-of-domain evaluation protocol on TuSimple, HSCM-Lane-S, HSCM-Lane-M, and HSCM-Lane-L achieve 27.2%, 27.6%, and 27.6% lane IoU, respectively, without target-domain fine-tuning. These results indicate that multi-level shifted-window context modeling improves the no-HSCM counterpart and yields competitive lane-mask segmentation results under these protocols. Future work will address temporal, instance-level, multimodal, planning, and safety validation.
Authors
- Tuong Le (ORCID: https://orcid.org/0000-0003-0909-4974)
- Thanh Nguyên Vũ
- Thanh Hien Vu
- Quang Toai Ton (ORCID: https://orcid.org/0009-0000-7958-6195)
Institutions
- Ho Chi Minh City University of Foreign Languages and Information Technology (VN)
- Ho Chi Minh City University of Industry and Trade (VN)
- HUTECH University
Publication Details
- Journal
- PLoS ONE
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1371/journal.pone.0359959
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00