acceptodds
Under review as a conference paper at ICLR 2027

BASIS: Evaluating Belief-Aware Solutions in Interactive Settings beyond Instructions

Abstract

Large language model (LLM) agents are increasingly capable of carrying out complex, long-horizon tasks, yet effective assistance requires more than faithfully executing user instructions. Users' requests reflect their current understanding of the problem, which may be incomplete, outdated, or incorrect. When this understanding diverges from the actual task state, literal instruction following can fail to accomplish the user's underlying goal. We introduce **BASIS** (***B**elief-**A**ware **S**olutions in **I**nteractive **S**ettings*), a benchmark that evaluates whether agents can infer the mental states underlying user requests, recognize belief-reality mismatches, and use this understanding to develop appropriate solutions. Rather than assessing mental-state inference solely through explicit questions, **BASIS** evaluates both the utility of such inferences and the quality of downstream solutions in user-agent interactions. Evaluation results show that current models often fail to account for users' mental states when generating solutions, and that solution quality is consistently higher when such information is explicitly elicited and made available during problem solving. We further investigate the practical value of mental-state consideration through comparisons with generic chain-of-thought prompting, controlled interventions on mental-state information, targeted training, and multi-agent collaboration experiments under partial observability. Together, these studies highlight the importance of not merely inferring others' mental states, but also using those understandings to provide grounded, goal-directed assistance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.