The Geometry of Ignorance: How LLMs Encode and Adjust a Bayesian Prior
Abstract
What does a language model predict when it has few clues? The answer lies in its unembedding geometry: a single direction of the unembedding matrix encodes the unigram distribution of the training corpus, which serves as the Bayesian prior the model falls back on when uncertain. This structure, termed the *direction of ignorance*, appears across all model families examined (Llama, Qwen, Gemma, and Pythia), ranging from 0.4B to 405B parameters. Projecting the final prediction state onto this direction yields a per-token *prior loading factor* , which declines steadily as the context becomes more informative. Formally, the same projection decomposes the prediction state into two orthogonal vectors interpretable as the two components of a tempered Bayesian update: a unigram prior raised to the exponent and a context-driven likelihood. This geometric-probabilistic connection makes meaningfully comparable across model sizes and families. Larger models generally exhibit lower prior reliance, and reasoning models exhibit lower reliance than their base counterparts. Finally, we show that serves as a causal, inference-time control: reducing it can promote context-driven continuations and improve Llama-3.2-3B-Instruct on GSM8K, whereas amplifying it or applying matched perturbations along random directions degrades performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.