A Conditioning Lens on Multi-Head Attention
Abstract
Transformers have reshaped machine learning in large part through multi-head attention, yet the optimization role of multiple heads remains less understood. In this work, we provide a theoretical and empirical perspective on multi-head attention through the conditioning of attention-layer Jacobians. Under explicit assumptions on the statistical structure of attention-head weights, we show that increasing attention heads can improve the conditioning of the attention Jacobian, suggesting a conditioning-like mechanism for multi-head attention. This analysis motivates the hypothesis that allocating capacity to additional heads can improve trainability more effectively than simply increasing depth. We test this hypothesis across a range of transformer architectures and domains, and find that head-rich configurations consistently exhibit better-conditioned attention Jacobians and improved optimization behaviour. Building on these findings, we redesign standard transformer models to use more heads and fewer layers, achieving reductions of 30–50% in parameter count, TFLOPs, and memory while preserving accuracy. These leaner architectures also converge in fewer training steps than commonly used transformer baselines. Overall, our results identify attention heads as an important architectural factor in transformer conditioning and provide a practical route to more efficient transformer design without degrading downstream performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.