Semantic-Conditioned Spectral Fusion for Label-Efficient Multispectral Adaptation of Frozen Remote-Sensing Vision–Language Models
Remote-sensing vision–language models (RS-VLMs) yield transferable, text-alignable representations, yet are pre-trained on three-band RGB imagery and therefore cannot exploit the red-edge, near-infrared and short-wave-infrared bands that carry much of the discriminative signal in multispectral data. Existing adaptations either restructure a frozen backbone's internal computation or pre-train a dedicated multispectral VLM on large image–caption corpora; both discard the component that distinguishes a VLM from a plain encoder – its text branch. We introduce Semantic-Conditioned Spectral Fusion (SCSF), in which the frozen model's zero-shot prediction against class text prototypes forms a per-sample semantic descriptor that modulates, through feature-wise linear modulation (FiLM), a lightweight trainable spectral branch prior to its fusion with the RGB feature; the model's own semantic estimate thus governs which spectral responses are emphasized. SCSF leaves the backbone frozen, trains only ~0.1M parameters, and is initialized as a no-op so it cannot degrade the backbone before training. On EuroSAT-MS and BigEarthNet-v2 it improves consistently over an unconditioned spectral-fusion baseline (up to +4.6% OA and +1.7% mAP), with gains concentrated in the low-label regime, and an ablation attributes the improvement to the semantic conditioning rather than to added capacity.
Authors
- Waleed Shoaib
- Bakhtawar Ismail
Institutions
- COMSATS University Islamabad (PK)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-12
- DOI
- https://doi.org/10.5281/zenodo.22729753
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- preprint