acceptodds
Under review as a conference paper at ICLR 2027

ELICITBench: Measuring Elicitation Capabilities of Interactive Agents

Abstract

Large Language Model (LLM) agents deployed in interactive settings often receive user requests to make recommendations that satisfy certain requirements. An agent can satisfy every stated requirement and still return something the user is ineligible for, because the fact that decides eligibility was never elicited. However, which question is worth asking cannot always be specified in the system prompt. Such tasks require LLM agents to be good at _eliciting_ the right requirements from the user. We study this _hidden-state elicitation_ problem with ELICITBench, in which a user withholds part of their constraints and preferences, and the agent explores a catalog of services, asks questions grounded in what it finds, and submits a ranked shortlist. Existing benchmarks score interaction with missing information that is acquirable through a deterministic interaction where questions can be inferred from the system instructions and tool arguments; here both the missing information and the path to acquiring it are created by exploration and elicitation. This task allows us to diagnose agent failure modes in interactive settings, attributing failures to a) not acquiring _enough_ information through search, b) asking wrong or sub-optimal questions, or c) asking fewer questions than the agent should have. Given perfect information, frontier models are largely capable of providing satisfactory replies, but most models drop by 40% when asked to elicit the right requirements. Furthermore, when we investigate how different deployment conditions (e.g., code harness, test-time compute, self-reflection) affect elicitation capabilities, we find they do not substantially help elicitation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.