The Shared Geometry of Language Models
Abstract
Language models with different sizes and architectures develop strikingly similar representation geometry, with specific observed arrangements such as cycles and helices often interpreted as traces of computation. We propose the simpler explanation that much of this shared geometry reflects the geometry at model input and output. We show that the shape of input embeddings extends into middle layers while representations converge to the geometry of their predictions well before the last layer. To understand what drives the geometry of these input and output weights, we train models without hidden layers or with only one attention layer and show that they recover a large fraction of the geometry shared across models. This suggests that low-order corpus statistics are a major source of the shared input and output geometry. We support this claim with tokenizer interventions that produce predictable geometric changes in both shallow and deeper models. Finally, we trace six previously reported structures in LLMs to the shape of the input or the shape of the answer. This suggests input and output geometry as an alternative explanation to rule out before attributing representation geometry to task-specific computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.