The Attribute Trap: Systematically Suboptimal Clarifying Questions After Tool-Call Retrieval
Abstract
Do task-oriented dialogue agents ground their clarifying questions in the records they have just retrieved? We study closed-set candidate disambiguation: a task-oriented dialogue agent retrieves multiple candidate records (orders, bookings, search results returned by a database query or an application programming interface (API) call) matching a user's request, and must ask a single clarifying question to identify which one of the candidates is meant. Because the candidate set is known, the expected information gain of every attribute the agent could ask about can be computed exactly, so an agent's question can be scored by its regret against the most informative one. This lets us separate genuine judgment failures from three confounds that distort naive failure rates: candidate sets that no natural attribute can resolve, questions that were in expectation as good a bet as the best one, and near-identifier attributes that are informative whatever the table contains. Our main evidence is interventional. On real candidate sets from the Schema-Guided Dialogue corpus (SGD) and the Olist e-commerce data we edit a single column and leave every other cell unchanged. The edits split nine open-weight instruction-tuned models (7–32B parameters) in two. When the column a model asks about most is only partly duplicated, so that it is still informative but no longer the best, four models, among them the two largest, stop asking it (3–21% of draws); four keep asking on 73–85%; one lies between. When the column is flattened so that it carries no information, six stop, two fall to about the chance rate, and one keeps asking at twice it. So four models do not react to the values in the table; four do, in different degrees. We call asking about an attribute that is not the most informative one an attribute trap. On unedited tables from four dataset families (SGD, the Multi-Domain Wizard-of-Oz corpus (MultiWOZ), the -bench family, Olist) every model falls into it; for the strongest models the loss is small and sits almost entirely on attributes such as an exact price, which are the most informative but which a real user may not know. The effect persists across prompt wordings, a recency-weighted prior and scoring thresholds, and it is not explained by counting errors alone: giving the agent the exact number of distinct values of every attribute does not reliably repair it, and stating the goal makes most models worse.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.