Attention-Sink Signatures Dissociate During Vision-Language Pretraining

Attention sinks are a widely studied target of interventions in language models and are often linked, though not causally, to hallucinations in vision-language models. A sink is defined by attention concentration (Sink₁ᵋ), but in text LMs it reliably arrives with two companions: a drained value norm (v-ratio) and massive activations (h-ratio) at the same position. We track all three separately during multimodal pretraining, without assuming they are one phenomenon. The model is a 222M vision-language model with a randomly initialized decoder, position 0 being an image token, no BOS, trained under four levers: softmax attention, output-gated softmax, unnormalized sigmoid attention, and initialization from a pretrained text LM. The three signatures dissociate: each training lever produces a different combination of them, consistently across 2-3 seeds per arm. On a 1B fresh-token run (2.39 effective visual epochs, single seed), the massive-activation proxy more than doubles while attention concentration stays at exactly zero at every threshold we test. Value-norm drain emerges as a third axis, beyond the massive-activation vs. sink dissociations known from text-only models.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23117576
Primary Topic
Mind wandering and attention
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Attention-Sink Signatures Dissociate During Vision-Language Pretraining

Samvat Tiwari
Zenodo (CERN European Organization for Nuclear Research)
Mind wandering and attention
preprint

Attention-Sink Signatures Dissociate During Vision-Language Pretraining

Samvat Tiwari
preprint en

Abstract

Attention sinks are a widely studied target of interventions in language models and are often linked, though not causally, to hallucinations in vision-language models. A sink is defined by attention concentration (Sink₁ᵋ), but in text LMs it reliably arrives with two companions: a drained value norm (v-ratio) and massive activations (h-ratio) at the same position. We track all three separately during multimodal pretraining, without assuming they are one phenomenon. The model is a 222M vision-language model with a randomly initialized decoder, position 0 being an image token, no BOS, trained under four levers: softmax attention, output-gated softmax, unnormalized sigmoid attention, and initialization from a pretrained text LM. The three signatures dissociate: each training lever produces a different combination of them, consistently across 2-3 seeds per arm. On a 1B fresh-token run (2.39 effective visual epochs, single seed), the massive-activation proxy more than doubles while attention concentration stays at exactly zero at every threshold we test. Value-norm drain emerges as a third axis, beyond the massive-activation vs. sink dissociations known from text-only models.

Zenodo (CERN European Organization for Nuclear Research)
Mind wandering and attention
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.