acceptodds
Under review as a conference paper at ICLR 2027

Measuring Internal Uncertainty in Language Model Agents: A Multi-Method Triangulation Approach

Abstract

Adaptive information acquisition requires an agent to assess uncertainty about its current judgement. For large language models (LLMs), however, it remains unclear which observable signals provide useful measures of this internal uncertainty. We use a controlled probabilistic reasoning task with evidence that is either ambiguous, strongly supports the correct answer, or strongly supports an incorrect answer. This design allows us to test whether candidate uncertainty measures respond to how conclusive the evidence is, even when it is misleading. We compare candidate measures derived from verbal reports, answer-option distributions, opt-out behaviour, reasoning traces, and internal representations, measured in GPT-OSS-20B and Qwen3.5-27B. We evaluate the ability of these candidate indicators to distinguish evidence conditions, retain a set of effective indicators, and assess their convergence across measurement methods. We construct two composite indices that consistently distinguish evidence conditions in both models. In Qwen3.5-27B, higher measured uncertainty is further associated with sampling beyond the model's currently preferred cue. These findings provide a measurement basis for investigating how internal uncertainty relates to information acquisition and decision-making.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.