acceptodds
Under review as a conference paper at ICLR 2027

Multi-Head Attention Residuals

Abstract

Transformers propagate information across depth through a single additive residual stream. Attention residuals relax this by letting each sublayer attend, through a learned softmax, over all earlier sublayer outputs - but that read shares one query across the entire width, forcing every feature subspace through the same depth distribution, at a cost that grows with model width. We introduce Multi-Head Attention Residuals (MHAR): reshape the routing query into H per-subspace heads, each with its own softmax over depth. The reshape adds zero parameters and negligible compute, and H=1 recovers attention residuals exactly. Trained from scratch on a STEM- and code-heavy corpus we construct, MHAR improves validation loss over a standard Transformer by 0.061/0.149/0.140 at 100M/350M/1B - the best of four methods in every setting - and the gains transfer to held-out perplexity and LAMBADA accuracy, survive a document-disjoint re-evaluation, and persist at a compute-optimal token budget rather than vanishing as an under-training artifact. Validation loss is U-shaped in H with a flat optimum at H=4-8; we adopt H=8. Temperature-matched controls attribute most of the gain to the flatter per-head softmax the reshape provides automatically - a single query's depth softmax over-sharpens as width grows - with a small residual consistent with per-subspace routing. Fused Triton kernels lift attention-residual training from 0.2-0.5x to 0.55-0.88x baseline throughput at near-baseline memory, and an identity-preserving conversion brings MHAR to 8B mid-training (+3.2 GSM8K, +3.1 GPQA).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.