acceptodds
Under review as a conference paper at ICLR 2027

Beyond Finding Facts: DISCOBench for Evaluating Clarification in Deep Search

Abstract

Search agents powered by large language models (LLMs) increasingly tackle complex information-seeking tasks through multi-step retrieval and reasoning. However, existing benchmarks often assume complete and explicit queries, leaving agents' ability to clarify underspecified requests insufficiently evaluated. In multi-step search, retrieved facts may support several plausible intermediate targets without revealing which one the user intends. Final-answer correctness alone offers limited insight into whether agents recognize and resolve these information gaps. We introduce DISCOBench, a benchmark for evaluating whether search agents recognize clarification needs, ask effective questions, and use feedback to advance through interdependent retrieval steps. Starting from verified multi-hop questions, we introduce ambiguous checkpoints with annotated discriminative clues and design a user simulator that provides these clues in response to clarification questions. The annotated checkpoints provide verifiable intermediate targets for assessing whether agents ask effective clarification questions and successfully advance after receiving feedback. DISCOBench contains 211 samples and 463 ambiguity instances across 11 knowledge domains and four ambiguity types. We evaluate task utility, ambiguity detection, clarification effectiveness, and interaction cost. Experiments on 12 LLMs show that even the best-performing model achieves only 63.7% task accuracy, highlighting the challenge of coordinating retrieval with proactive clarification in multi-step search.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.