A Normal Form for Softmax Attention
Abstract
We show that every softmax attention head admits a unique three-term normal form with interpretable components. An inspired choice of gauge in the score function clearly separates geometric and topological contributions, including a metric term which is the complement of a Laplacian in a head-induced geometry, a 1-cochain whose circulation encodes the area of token triangles, and a rank-one offset that biases the keys independently of the queries. This enables a precise characterization of when softmax attention is purely diffusive and when differently represented discrete Laplacian operators are the same. Applying the proposed normal form shows that softmax attention generically implements a quasilinear advection-diffusion equation satisfying a discrete maximum principle, guaranteeing stable information propagation regardless of its weights. Additional analysis shows that the metric and circulation terms in the normal form determine a guaranteed rate at which a softmax head mixes tokens, while the remaining offset term bounds the amount of attention that each key can receive. Ultimately, this leads to a reinterpretation of observed phenomena related to attention sinks and oversmoothing that may be useful for understanding deep attention networks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.