acceptodds
Under review as a conference paper at ICLR 2027

How Do Language Models Manage Uncertainty?: A Case Study of the Number Game

Abstract

Language models must navigate ambiguity to successfully complete sentences or perform tasks; however, we have little mechanistic understanding of how they represent or handle such uncertainty. To understand this better, we study how Llama-3.1 (8B) plays the Number Game, a concept learning task in which participants view a set of numbers that conform to an unknown concept, and must determine which other numbers conform. Llama's behavior, like humans', appears Bayesian; we search for the mechanisms underlying this seemingly normative behavior. We first find that Llama's representational geometry constrains the concepts that it considers as hypotheses, allowing it to consider only a specific range of possible hypotheses. We next find that Llama has internal circuitry that identifies common concepts in its input, and downweights poorly supported ones. Finally, we find that Llama's increased confidence upon seeing more inputs stems largely from an attention mechanism that simply tracks the number of examples seen. We use these insights to design a new model that better predicts Llama's behavior, including on new concept types. In sum, we offer a detailed account of how Llama manages its uncertainty and what mechanisms underlie this capacity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.