Beyond Finding Facts: DISCOBench for Evaluating Clarification in Deep Search
Abstract
Search agents powered by large language models (LLMs) increasingly tackle complex information-seeking tasks through multi-step retrieval and reasoning. However, existing benchmarks often assume complete and explicit queries, leaving agents' ability to clarify underspecified requests insufficiently evaluated. In multi-step search, retrieved facts may support several plausible intermediate targets without revealing which one the user intends. Final-answer correctness alone offers limited insight into whether agents recognize and resolve these information gaps. We introduce DISCOBench, a benchmark for evaluating whether search agents recognize clarification needs, ask effective questions, and use feedback to advance through interdependent retrieval steps. Starting from verified multi-hop questions, we introduce ambiguous checkpoints with annotated discriminative clues and design a user simulator that provides these clues in response to clarification questions. The annotated checkpoints provide verifiable intermediate targets for assessing whether agents ask effective clarification questions and successfully advance after receiving feedback. DISCOBench contains 211 samples and 463 ambiguity instances across 11 knowledge domains and four ambiguity types. We evaluate task utility, ambiguity detection, clarification effectiveness, and interaction cost. Experiments on 12 LLMs show that even the best-performing model achieves only 63.7% task accuracy, highlighting the challenge of coordinating retrieval with proactive clarification in multi-step search.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.