Why Language Models Disambiguate Contexts Despite Similar Embeddings
Abstract
Language models present an apparent paradox where the representations of a token across distinct contexts exhibit very high cosine similarity, and yet they cleanly disambiguate contexts. Here, we analyze the tangent space dynamics of the embedding and unembedding spaces, to understand the influence of context and layer depth on the output. We introduce a two-fold decomposition that separates representations into (i) a shared component common across contexts and its context-specific orthogonal complement, and (ii) readout-sensitive and readout-insensitive subspaces determined by their direct linear projection through the unembedding matrix. From large scale empirical studies on 16 models of 160M to 9B parameters, we observe that the shared context-invariant subspace is high-dimensional, occupying a large fraction of the ambient embedding space. The readout is dominated by the component shared across contexts, which is why representations across contexts look nearly identical. Separation survives because the shared components cancel and the contrast component is enriched tenfold in the subspace the readout acts on.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.