Attention Mean Fields Predict Average Representation Dynamics and Illuminate Early Training
Abstract
A language model's representation geometry is not predetermined; it evolves as computation proceeds. Each layer reshapes geometry in part through the effects of attention, which updates each token's representation with information drawn from its context. A faithful account of a model's geometry must capture that dynamic process. To that end, we introduce a mean-field analysis of attention. The average attention a head pays from one token to another defines a kernel that carries representations forward through the model and describes how geometry is transformed. Conditioned on a substantial corpus, the kernel predicts, from the input embeddings and frozen weights alone, the trajectory of each token's mean representation. These predictions are accurate across models spanning GPT-2 to Qwen-3-14B. We show that model-oblivious alternatives, including one built from token co-occurrence, are all poorer predictors. We then develop an attention mean field with more fine-grained conditioning, and demonstrate its value. Specifically, we show that early in training, on real text, replacing every attention head's output with its context-conditional mean field leaves the model's loss unchanged. This effect occurs across three scales of Pythia as well as in OLMo-2-7B. As training progresses through the onset of induction, model behavior and mean-field behavior diverge, with representations becoming contextualized and attention becoming correlated with transported values.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.