Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View
Semantic segmentation for autonomous driving is challenged by limited field-of-view (FoV), where occlusions and restricted viewpoints reduce available visual information. Existing approaches often incur considerable computational overhead, which is unsuitable for autonomous driving systems. To build robust and lightweight semantic segmentation models, attention-based knowledge distillation (KD) can be an efficient alternative. Nevertheless, KD remains primarily visual and provides limited information when visual cues are incomplete or ambiguous. In contrast, biological visual systems integrate visual cues with contextual and semantic knowledge to support robust scene perception. Motivated by this biomimetic principle, language-guided methods provide high-level semantic knowledge, offering another promising direction for enhancing visual understanding. However, the effectiveness of attention-based KD under limited FoV conditions and its interaction with semantic knowledge remain largely unexplored. To systematically investigate this problem, we devise Semantic-guided Attentive Feature Distillation (SAFD), a KD-based framework that combines attention-based feature distillation with semantic guidance derived from large language models (LLMs) without introducing additional inference-time computation. Through comprehensive analyses of multiple distillation strategies and textual representations using the framework, we observe that semantic guidance generally provides additional benefits when the gains from visual distillation are limited, while its effectiveness varies across different attention-based distillation strategies and textual representations. Furthermore, our analyses reveal that different attention mechanisms exploit semantic guidance in distinct ways, with consistent benefits across different pretrained text encoders and improved generalization. These findings provide practical insights into integrating linguistic semantics with visual knowledge transfer under limited FoV.
Authors
- Eun Som Jeon (ORCID: https://orcid.org/0000-0002-1112-4653)
- Jin Hyeok Ryu
- Seung Woo
Institutions
- Seoul National University of Science and Technology (KR)
Publication Details
- Journal
- Biomimetics
- Published
- 2026-09-16
- DOI
- https://doi.org/10.3390/biomimetics11090665
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00