A Hidden Decoding Risk Makes Bigger Language Models Less Reliable
Abstract
Bigger language models are less reliable. Across three families, three benchmarks and 11 models from 0.6B to 32B, including in-the-wild chat logs, scaling closes the start-of-response knowledge gap up to 7x while within-response knowledge degradation grows up to 39x. We trace that residual to one variable, the per-position disagreement delta = log p_M - log p_O against a stronger oracle, whose second moment splits exactly into bias^2 KL(p_M || p_O)^2 and decoding risk Var[delta]. That split is an interpretability statement before it is a statistical one: the model's self-readable uncertainty H(p_M) enters only the bias term, so the risk term has no model-readable component. Risk also takes a growing share of the squared error with scale, 31% to 49% from 1.7B to 14B. At a fabrication H(p_M) relaxes within one token while risk persists for 3.4-4.6 tokens, leaving a confident-but-precarious regime that bridges consecutive fabrications (+69% at 14B). Contracting that risk at fixed KL removes 35-74% of web-verified hallucinations across six models and three families. Semantic entropy fires less on that branch (p<10^-3) though it carries nearly 4x the fabrications. Bigger models snowball mistakes faster, through a failure mode that is dominant, self-perpetuating, causal and invisible to the model itself.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.