Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View

Semantic segmentation for autonomous driving is challenged by limited field-of-view (FoV), where occlusions and restricted viewpoints reduce available visual information. Existing approaches often incur considerable computational overhead, which is unsuitable for autonomous driving systems. To build robust and lightweight semantic segmentation models, attention-based knowledge distillation (KD) can be an efficient alternative. Nevertheless, KD remains primarily visual and provides limited information when visual cues are incomplete or ambiguous. In contrast, biological visual systems integrate visual cues with contextual and semantic knowledge to support robust scene perception. Motivated by this biomimetic principle, language-guided methods provide high-level semantic knowledge, offering another promising direction for enhancing visual understanding. However, the effectiveness of attention-based KD under limited FoV conditions and its interaction with semantic knowledge remain largely unexplored. To systematically investigate this problem, we devise Semantic-guided Attentive Feature Distillation (SAFD), a KD-based framework that combines attention-based feature distillation with semantic guidance derived from large language models (LLMs) without introducing additional inference-time computation. Through comprehensive analyses of multiple distillation strategies and textual representations using the framework, we observe that semantic guidance generally provides additional benefits when the gains from visual distillation are limited, while its effectiveness varies across different attention-based distillation strategies and textual representations. Furthermore, our analyses reveal that different attention mechanisms exploit semantic guidance in distinct ways, with consistent benefits across different pretrained text encoders and improved generalization. These findings provide practical insights into integrating linguistic semantics with visual knowledge transfer under limited FoV.

Authors

Institutions

Publication Details

Journal
Biomimetics
Published
2026-09-16
DOI
https://doi.org/10.3390/biomimetics11090665
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View

Eun Som Jeon, Jin Hyeok Ryu, Seung Woo
Biomimetics
Multimodal Machine Learning Applications
article

Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View

Eun Som Jeon, Jin Hyeok Ryu, Seung Woo
article en

Abstract

Semantic segmentation for autonomous driving is challenged by limited field-of-view (FoV), where occlusions and restricted viewpoints reduce available visual information. Existing approaches often incur considerable computational overhead, which is unsuitable for autonomous driving systems. To build robust and lightweight semantic segmentation models, attention-based knowledge distillation (KD) can be an efficient alternative. Nevertheless, KD remains primarily visual and provides limited information when visual cues are incomplete or ambiguous. In contrast, biological visual systems integrate visual cues with contextual and semantic knowledge to support robust scene perception. Motivated by this biomimetic principle, language-guided methods provide high-level semantic knowledge, offering another promising direction for enhancing visual understanding. However, the effectiveness of attention-based KD under limited FoV conditions and its interaction with semantic knowledge remain largely unexplored. To systematically investigate this problem, we devise Semantic-guided Attentive Feature Distillation (SAFD), a KD-based framework that combines attention-based feature distillation with semantic guidance derived from large language models (LLMs) without introducing additional inference-time computation. Through comprehensive analyses of multiple distillation strategies and textual representations using the framework, we observe that semantic guidance generally provides additional benefits when the gains from visual distillation are limited, while its effectiveness varies across different attention-based distillation strategies and textual representations. Furthermore, our analyses reveal that different attention mechanisms exploit semantic guidance in distinct ways, with consistent benefits across different pretrained text encoders and improved generalization. These findings provide practical insights into integrating linguistic semantics with visual knowledge transfer under limited FoV.

BiomimeticsVol. 11(9)
Seoul National University of Science and Technology (KR)
Quality Education
Openalex Percentile: Top 13%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View — Eun Som Jeon, Jin Hyeok Ryu, et al. · Biomimetics (2026) | TGRS Research Map | TGRS