acceptodds
Under review as a conference paper at ICLR 2027

When Is Softmax Attention a Gradient? Exact Characterization and Conditioning

Abstract

Energy and potential representations provide structural explanations for the fixed points and dynamics of tied or specially structured attention maps. Fixed-memory softmax attention can be a gradient even when its keys and effective values are untied, but this depends on the additive geometry of the keys. For , we give a finite necessary-and-sufficient criterion for a global potential by grouping equal key-pair sums. Generic memories force a common scaling and translation of the keys; binary Cartesian memories admit exactly independent coordinate scalings. We characterize constant-metric convex representations and derive two-sided finite-query bounds on the distance to the conservative value space from measured Jacobian asymmetry, with explicit conditioning and observation-error dependence. A near-Cartesian family shows why conditioning is essential: curl tends to zero while both the parameter distance and the best field approximation within the same-key conservative class remain bounded away from zero. Deterministic numerical checks validate the exact characterization and finite-query certificates, and illustrate the predicted loss of conditioning near pair-sum collisions. Finally, an analytic bifurcation construction shows that arbitrarily small incompatible perturbations can produce oscillation near a reference map's stability boundary. These results distinguish exact compatibility, reliable approximate certification, and dynamical stability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.