Rojalina Priyadarshini, Subham Divakar · Transactions on Artificial Intelligence 2026 · 2026
DOI: 10.53941/tai.2026.100014
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Standard Multi-Head Attention (MHA) computes contextual representations in a single forward pass, providing no mechanism to detect or correct diffuse or suboptimal attention distributions. Existing efforts to improve attention have primarily targeted computational efficiency or sequence length, leaving attention quality itself largely unaddressed. We propose IMHA (Entropy-Guided Iterative Multi-Head Attention), a within-layer attention formulation that progressively refines token representations over T iterations using shared projection weights. At each step, the Shannon entropy of each token’s per-head attention row—normalised against the position-dependent maximum to remove a structural artefact of causal masking—serves as an intrinsic quality signal: tokens with peaked (low-entropy) attention are amplified while those with diffuse (high-entropy) attention are suppressed. A gated residual update incorporating a prefix-causal global context vector ensures stable convergence. The formulation is strictly causal, verified at the gradient level across all ablation configurations. We prove (Proposition 1) that under mild assumptions the tokenspecific content of an attention output decays monotonically with attention diffuseness, providing principled motivation for entropy weighting. We further show (Proposition 2) that if iterative refinement sharpens the attention distribution, discriminative content is non-decreasing across iterations. We evaluate across three model scales (Small 14 M, Medium 45 M, Large 87 M parameters) on WikiText-2, Penn Treebank, and TinyStories, with downstream transfer to SST-2 sentiment classification and AG News topic classification. At medium scale, IMHA reduces WikiText-2 perplexity from 47.30 ± 0.18 to 44.85 ± 0.16 over the MHA baseline (paired t-test, p = 3.2 × 10−6, d = 11.4, after Holm–Bonferroni correction over 18 comparisons), outperforms parameter-matched capacity baselines by 1.40 points, and surpasses attention-quality methods including α-entmax, RealFormer, Universal Transformer (FLOPs-matched), and entropy-regularised MHA. Critically, IMHA outperforms entropy-regularised MHA by 1.18 points, demonstrating that the internal, within-layer delivery of the entropy signal—rather than entropy as an external loss term—is what drives the improvement. The relative gain is stable across scales (4.8%, 4.9%, 4.6% at Small, Medium, and Large respectively), and the pretrained representations transfer to classification: IMHA achieves 93.2 ± 0.3% on SST-2 vs. 91.8 ± 0.4% for MHA, and 90.1 ± 0.4% vs. 88.7 ± 0.5% on AG News. Total overhead is a uniform ∼8% across parameters, FLOPs, training time, and inference latency.
No comments yet — start the discussion below.