Affine Anticipation: LLM Answer Representations Are Largely Linearly Predictable from Context
Abstract
We show that the representation of an LLM's answer is largely a linear function of the representation of its context. We fit metamodels from the residual-stream state at the last context token (the context vector) to the mean answer-token activation (the answer vector), which we call context-answer metamodels. On half a million real-chat pairs from Qwen2.5-7B-Instruct, a linear metamodel reaches a held-out of 0.81, against 0.86 for a nonlinear MLP. The linear metamodel predicts high-level and identity-related aspects of the answer better than low-level and topic-related aspects. The relationship is already present in the base model, is reshaped by supervised fine-tuning (SFT), and is largely preserved by direct preference optimization (DPO) and reinforcement learning with verifiable rewards (RLVR). The state at the end of a chain-of-thought (CoT) trace predicts the answer better than the context state does, and predictability tracks model capability across ten models. The metamodel can also help forecast unwanted behaviors before generation and retrieve contexts likely to elicit them. More broadly, our results add to the evidence for linear structure in transformer representations and its link to language modelling capabilities, while motivating further study into context-answer metamodels as simple and useful approximations of LLM behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.