An Infinite Information Exponent that Gradient Descent Still Learns at the Optimal Rate: The Value Catalyst in Softmax Attention
Abstract
The information exponent of a single-index target describes how gradient descent learns it: an exponent- target can trap online SGD for samples, while an infinite exponent ordinarily indicates a direction that gradient descent cannot recover. We show that this interpretation fails for a coupled attention model. Consider a softmax head that selects a token through a key direction and reads a scalar through an orthogonal value direction , The selection direction has infinite information exponent: the label is mean-zero conditional on the key scores, so every key-only correlational query and every single-index surrogate gradient vanishes. Nevertheless, projected gradient descent on the invariant slice of the coupled key-value loss learns with samples per step; unrestricted random initializations exhibit the same cascade numerically. The value direction has information exponent one; once it is partially learned, the key gradient gains the explicit component where and is the mean softmax collision probability. This value catalyst produces a value-to-selection cascade with dimension-free population convergence. Squaring the label exposes the same selection direction at generative exponent two and yields a one-step spectral estimator with a linear dimensional exponent up to logarithmic factors. For every noise level, an oracle lower bound that reveals all exact key scores still requires samples for constant projective accuracy; at fixed positive noise, a Fano bound additionally gives explicit parameter dependence. The phenomenon extends to additive and multiplicative combinations of orthogonal-subspace heads, and CPU experiments check the predicted gradients, spectra, and scaling laws.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.