The Dimensional Cost of Context in Language Models
Abstract
Low-rank structure in language-model activations is often taken as evidence that the residual stream can be narrowed. However, a subspace that works for one prediction may not work for another. We define the context requirement dimension as the smallest rank of a (linear) map that preserves the model's output distributions within a fixed tolerance. We show that a prediction may depend on few directions, but these directions vary across contexts and together cover most of the residual width. At the output, preserving the distributions within our tolerance requires roughly 67–83% of the residual width across 21 models from nine families. The requirement is lower inside the network but generally increases towards the output. During generation, narrowing the stream shifts output towards frequent tokens. Therefore, local low dimensionality gives a misleading estimate of the width needed across a corpus.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.