Softmax Attention as a Covariance Readout: A Unified View of In-Context Learning and Repetition
Abstract
Large language models (LLMs) exhibit two seemingly unrelated behaviors: in-context learning (ICL) and repetitive generation. In both cases, the model appears to summarize the context into a population-level statistic to guide subsequent generation while discarding token-level detail. We investigate whether this "summarization and forgetting" can be derived from the attention mechanism itself. For stationary inputs with elliptical marginals, the softmax attention outputs converge almost surely to in the long-context regime, where and denote the input mean and covariance, respectively, and is a scalar. Beyond the mean contribution, this limit provides a readout of the input’s second-order statistics. Its dependence on the current input token alone leads to two consequences. (i) For in-context linear regression tasks, a residual softmax attention layer can emulate one gradient-descent step. (ii) When propagated across multiple transformer layers, this readout drives the terminal hidden state toward a deterministic function of the current token. Autoregressive generation thus reduces asymptotically to a first-order Markov chain, whose periodic orbits provide a structural explanation for repetitive generation. These two phenomena emerge as facets of a common covariance-readout principle.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.