Attention-Sink Signatures Dissociate During Vision-Language Pretraining
Attention sinks are a widely studied target of interventions in language models and are often linked, though not causally, to hallucinations in vision-language models. A sink is defined by attention concentration (Sink₁ᵋ), but in text LMs it reliably arrives with two companions: a drained value norm (v-ratio) and massive activations (h-ratio) at the same position. We track all three separately during multimodal pretraining, without assuming they are one phenomenon. The model is a 222M vision-language model with a randomly initialized decoder, position 0 being an image token, no BOS, trained under four levers: softmax attention, output-gated softmax, unnormalized sigmoid attention, and initialization from a pretrained text LM. The three signatures dissociate: each training lever produces a different combination of them, consistently across 2-3 seeds per arm. On a 1B fresh-token run (2.39 effective visual epochs, single seed), the massive-activation proxy more than doubles while attention concentration stays at exactly zero at every threshold we test. Value-norm drain emerges as a third axis, beyond the massive-activation vs. sink dissociations known from text-only models.
Authors
- Samvat Tiwari (ORCID: https://orcid.org/0009-0001-2636-0544)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23117576
- Primary Topic
- Mind wandering and attention
- Type
- preprint