Rethinking Attention: Entropy-Guided Iterative Refinement for Language Models

Standard Multi-Head Attention (MHA) computes contextual representations in a single forward pass, providing no mechanism to detect or correct diffuse or suboptimal attention distributions. Existing efforts to improve attention have primarily targeted computational efficiency or sequence length, leaving attention quality itself largely unaddressed. We propose IMHA (Entropy-Guided Iterative Multi-Head Attention), a within-layer attention formulation that progressively refines token representations over T iterations using shared projection weights. At each step, the Shannon entropy of each token’s per-head attention row—normalised against the position-dependent maximum to remove a structural artefact of causal masking—serves as an intrinsic quality signal: tokens with peaked (low-entropy) attention are amplified while those with diffuse (high-entropy) attention are suppressed. A gated residual update incorporating a prefix-causal global context vector ensures stable convergence. The formulation is strictly causal, verified at the gradient level across all ablation configurations. We prove (Proposition 1) that under mild assumptions the tokenspecific content of an attention output decays monotonically with attention diffuseness, providing principled motivation for entropy weighting. We further show (Proposition 2) that if iterative refinement sharpens the attention distribution, discriminative content is non-decreasing across iterations. We evaluate across three model scales (Small 14 M, Medium 45 M, Large 87 M parameters) on WikiText-2, Penn Treebank, and TinyStories, with downstream transfer to SST-2 sentiment classification and AG News topic classification. At medium scale, IMHA reduces WikiText-2 perplexity from 47.30 ± 0.18 to 44.85 ± 0.16 over the MHA baseline (paired t-test, p = 3.2 × 10−6, d = 11.4, after Holm–Bonferroni correction over 18 comparisons), outperforms parameter-matched capacity baselines by 1.40 points, and surpasses attention-quality methods including α-entmax, RealFormer, Universal Transformer (FLOPs-matched), and entropy-regularised MHA. Critically, IMHA outperforms entropy-regularised MHA by 1.18 points, demonstrating that the internal, within-layer delivery of the entropy signal—rather than entropy as an external loss term—is what drives the improvement. The relative gain is stable across scales (4.8%, 4.9%, 4.6% at Small, Medium, and Large respectively), and the pretrained representations transfer to classification: IMHA achieves 93.2 ± 0.3% on SST-2 vs. 91.8 ± 0.4% for MHA, and 90.1 ± 0.4% vs. 88.7 ± 0.5% on AG News. Total overhead is a uniform ∼8% across parameters, FLOPs, training time, and inference latency.

Authors

Institutions

Publication Details

Journal
Transactions on Artificial Intelligence
Published
2026-09-29
DOI
https://doi.org/10.53941/tai.2026.100014
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Rethinking Attention: Entropy-Guided Iterative Refinement for Language Models

Subham Divakar, Rojalina Priyadarshini
Transactions on Artificial Intelligence
Topic Modeling
article

Rethinking Attention: Entropy-Guided Iterative Refinement for Language Models

Subham Divakar, Rojalina Priyadarshini
article en

Abstract

Standard Multi-Head Attention (MHA) computes contextual representations in a single forward pass, providing no mechanism to detect or correct diffuse or suboptimal attention distributions. Existing efforts to improve attention have primarily targeted computational efficiency or sequence length, leaving attention quality itself largely unaddressed. We propose IMHA (Entropy-Guided Iterative Multi-Head Attention), a within-layer attention formulation that progressively refines token representations over T iterations using shared projection weights. At each step, the Shannon entropy of each token’s per-head attention row—normalised against the position-dependent maximum to remove a structural artefact of causal masking—serves as an intrinsic quality signal: tokens with peaked (low-entropy) attention are amplified while those with diffuse (high-entropy) attention are suppressed. A gated residual update incorporating a prefix-causal global context vector ensures stable convergence. The formulation is strictly causal, verified at the gradient level across all ablation configurations. We prove (Proposition 1) that under mild assumptions the tokenspecific content of an attention output decays monotonically with attention diffuseness, providing principled motivation for entropy weighting. We further show (Proposition 2) that if iterative refinement sharpens the attention distribution, discriminative content is non-decreasing across iterations. We evaluate across three model scales (Small 14 M, Medium 45 M, Large 87 M parameters) on WikiText-2, Penn Treebank, and TinyStories, with downstream transfer to SST-2 sentiment classification and AG News topic classification. At medium scale, IMHA reduces WikiText-2 perplexity from 47.30 ± 0.18 to 44.85 ± 0.16 over the MHA baseline (paired t-test, p = 3.2 × 10−6, d = 11.4, after Holm–Bonferroni correction over 18 comparisons), outperforms parameter-matched capacity baselines by 1.40 points, and surpasses attention-quality methods including α-entmax, RealFormer, Universal Transformer (FLOPs-matched), and entropy-regularised MHA. Critically, IMHA outperforms entropy-regularised MHA by 1.18 points, demonstrating that the internal, within-layer delivery of the entropy signal—rather than entropy as an external loss term—is what drives the improvement. The relative gain is stable across scales (4.8%, 4.9%, 4.6% at Small, Medium, and Large respectively), and the pretrained representations transfer to classification: IMHA achieves 93.2 ± 0.3% on SST-2 vs. 91.8 ± 0.4% for MHA, and 90.1 ± 0.4% vs. 88.7 ± 0.5% on AG News. Total overhead is a uniform ∼8% across parameters, FLOPs, training time, and inference latency.

Transactions on Artificial IntelligenceVol. 2(1)
C.V. Raman Global University
Reduced inequalities
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.