A Geometry-Dependent Theory of Softmax-Attention In-Context Learning Dynamics
Abstract
Transformers learn new tasks from examples placed in the prompt, yet theory of how gradient descent trains softmax attention is largely limited to linear tasks and orthogonal features. We study one-layer softmax attention trained on in-context regression with nonlinear task functions over non-orthogonal feature prototypes. Our analysis rests on an exact decomposition of the population loss and gradient: the task distribution enters only through the covariance of its outputs across prototypes, and the representation only through its Gram matrix, which transports every gradient from both sides. Task contrast therefore sets the clock of training, whereas representation geometry decides whether and how fast training converges. For every full-rank geometry we prove monotone descent and an inverse-accuracy rate once attention has found the right prototype; a geometry that violates self-similarity can trap training at nonzero loss, whereas small correlations and transitive geometries such as prototypes on a ring converge globally from zero initialization. Under symmetric geometry, training reduces exactly to a scalar recurrence with two laws: coarse routing, whose duration grows quadratically in the number of prototypes, and inverse-square refinement. The ratio of task contrast to target accuracy decides which law is active when the target is met, giving two exact thresholds, a crossover at which refinement dominates training time, and a learning-rate policy that decides whether stronger contrast speeds up or slows down convergence. Exact integrations and finite-sample SGD on the token-level model confirm the predicted boundaries and scaling laws.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.