Softmax Attention is an Exact Gradient Step: Predictable Finite Responses from RoPE Derivatives
Abstract
Transformers are said to learn in context by gradient descent, but the claim is proven only where pretrained models do not operate: in linear attention, or at weights constructed to make it true. For the softmax attention they use, it survives only as an approximation that linearises the exponential, and whether pretrained models implement gradient descent at all is disputed (Shen et al., 2024). We show that no approximation is needed: every softmax attention head, for any weights, context and query, computes exactly one step of gradient descent. Its output is the uniform average of its values plus the prediction, at its query, of a linear model trained by one gradient step of weighted least-squares regression from keys to centred values. Each example is weighted by the exponential’s divided difference at its attention score, so every weight is positive and rises with the score. As scores shrink, the weights become uniform and the step becomes the linear-attention gradient step of von Oswald et al. (2023). The loss’s curvature vanishes exactly on query edits that leave attention unchanged, and one linear solve finds the smallest edit realising requested attention ratios. The same identity, taken between two states of the head, gives the exact response to any finite edit, including moving a token, whose rotation follows from integrating RoPE’s derivative. In pretrained Qwen2.5 and SmolLM2, the gradient step reproduces every evaluated head to floating-point precision, and across 153,600 executed positional edits, exact responses cut the error of linearised predictions by 73.6–98.3%. Softmax attention does not approximate gradient descent in context. It performs it, exactly, in every head.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.