Semantic-Conditioned Spectral Fusion for Label-Efficient Multispectral Adaptation of Frozen Remote-Sensing Vision–Language Models

Remote-sensing vision–language models (RS-VLMs) yield transferable, text-alignable representations, yet are pre-trained on three-band RGB imagery and therefore cannot exploit the red-edge, near-infrared and short-wave-infrared bands that carry much of the discriminative signal in multispectral data. Existing adaptations either restructure a frozen backbone's internal computation or pre-train a dedicated multispectral VLM on large image–caption corpora; both discard the component that distinguishes a VLM from a plain encoder – its text branch. We introduce Semantic-Conditioned Spectral Fusion (SCSF), in which the frozen model's zero-shot prediction against class text prototypes forms a per-sample semantic descriptor that modulates, through feature-wise linear modulation (FiLM), a lightweight trainable spectral branch prior to its fusion with the RGB feature; the model's own semantic estimate thus governs which spectral responses are emphasized. SCSF leaves the backbone frozen, trains only ~0.1M parameters, and is initialized as a no-op so it cannot degrade the backbone before training. On EuroSAT-MS and BigEarthNet-v2 it improves consistently over an unconditioned spectral-fusion baseline (up to +4.6% OA and +1.7% mAP), with gains concentrated in the low-label regime, and an ablation attributes the improvement to the semantic conditioning rather than to added capacity.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-12
DOI
https://doi.org/10.5281/zenodo.22729753
Primary Topic
Multimodal Machine Learning Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Semantic-Conditioned Spectral Fusion for Label-Efficient Multispectral Adaptation of Frozen Remote-Sensing Vision–Language Models

Waleed Shoaib, Bakhtawar Ismail
Zenodo (CERN European Organization for Nuclear Research)
Multimodal Machine Learning Applications
preprint

Semantic-Conditioned Spectral Fusion for Label-Efficient Multispectral Adaptation of Frozen Remote-Sensing Vision–Language Models

Waleed Shoaib, Bakhtawar Ismail
preprint en

Abstract

Remote-sensing vision–language models (RS-VLMs) yield transferable, text-alignable representations, yet are pre-trained on three-band RGB imagery and therefore cannot exploit the red-edge, near-infrared and short-wave-infrared bands that carry much of the discriminative signal in multispectral data. Existing adaptations either restructure a frozen backbone's internal computation or pre-train a dedicated multispectral VLM on large image–caption corpora; both discard the component that distinguishes a VLM from a plain encoder – its text branch. We introduce Semantic-Conditioned Spectral Fusion (SCSF), in which the frozen model's zero-shot prediction against class text prototypes forms a per-sample semantic descriptor that modulates, through feature-wise linear modulation (FiLM), a lightweight trainable spectral branch prior to its fusion with the RGB feature; the model's own semantic estimate thus governs which spectral responses are emphasized. SCSF leaves the backbone frozen, trains only ~0.1M parameters, and is initialized as a no-op so it cannot degrade the backbone before training. On EuroSAT-MS and BigEarthNet-v2 it improves consistently over an unconditioned spectral-fusion baseline (up to +4.6% OA and +1.7% mAP), with gains concentrated in the low-label regime, and an ablation attributes the improvement to the semantic conditioning rather than to added capacity.

Zenodo (CERN European Organization for Nuclear Research)
COMSATS University Islamabad (PK)
Reduced inequalities
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Semantic-Conditioned Spectral Fusion for Label-Efficient Multispectral Adaptation of Frozen Remote-Sensing Vision–Language Models — Waleed Shoaib, Bakhtawar Ismail · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS