Measuring Internal Uncertainty in Language Model Agents: A Multi-Method Triangulation Approach
Abstract
Adaptive information acquisition requires an agent to assess uncertainty about its current judgement. For large language models (LLMs), however, it remains unclear which observable signals provide useful measures of this internal uncertainty. We use a controlled probabilistic reasoning task with evidence that is either ambiguous, strongly supports the correct answer, or strongly supports an incorrect answer. This design allows us to test whether candidate uncertainty measures respond to how conclusive the evidence is, even when it is misleading. We compare candidate measures derived from verbal reports, answer-option distributions, opt-out behaviour, reasoning traces, and internal representations, measured in GPT-OSS-20B and Qwen3.5-27B. We evaluate the ability of these candidate indicators to distinguish evidence conditions, retain a set of effective indicators, and assess their convergence across measurement methods. We construct two composite indices that consistently distinguish evidence conditions in both models. In Qwen3.5-27B, higher measured uncertainty is further associated with sampling beyond the model's currently preferred cue. These findings provide a measurement basis for investigating how internal uncertainty relates to information acquisition and decision-making.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.