acceptodds
Under review as a conference paper at ICLR 2027

Eliciting LLM Propositional Beliefs in Agentic Tasks: Consistency, Accuracy, and Response Tendencies

Abstract

Large language models (LLMs) are increasingly deployed as autonomous agents in interactive environments that require multi-step planning, reasoning, and state tracking. Evaluating an agent's beliefs about environment states and task progress can therefore help diagnose behavioral failures. However, reference belief labels are often unavailable or difficult to obtain. In this paper, we introduce a general, LLM-assisted evaluation pipeline that combines proposition generation with reference labels. The pipeline grounds LLM-authored proposition templates in agent trajectories and obtains labels from an LLM-assisted reference observer using formal logical rules and Bayesian inference. Using this pipeline, we systematically evaluate models' self-reported belief consistency and accuracy across World, Epistemic, Affordance, and Decision propositions in ALFWorld, WebShop, and Noisy Cabinets and examine their relationship with task success. T/F-only Affordance and Decision accuracy are positively associated with task success after centering within environments. Furthermore, we find that response tendencies help explain both observed consistency and accuracy variation: the tendency to answer Unknown contributes to cross-model differences in consistency, and a True/False tendency model explains degradation in accuracy and negation consistency as context grows.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.