acceptodds
Under review as a conference paper at ICLR 2027

DEEPSEARCHDAG: A DAG-STRUCTURED BENCHMARK FOR NODE-LEVEL EVALUATION OF DEEP SEARCH AGENTS

Abstract

Existing benchmarks for evaluating search agents predominantly verify only the final answer while treating intermediate sub-problems as a black box, providing no structural signal for distinguishing well-supported coverage from a potential lucky hit. Moreover, they typically involve a single target entity or a homogeneous set, lacking the multi-entity heterogeneity of real-world search tasks; benchmarks that jointly stress both search depth and width remain scarce. We present DeepSearchDAG (DSD), a DAG-structured benchmark built on a real-world online business dataset, decomposing queries involving multiple heterogeneous entities into atomic sub-problems connected by dependency edges, enabling node-level evaluation of both depth and width. DSD places representative search-agent benchmarks in a common node-level view and extends it to checked DAG dependencies. We propose NodeF1 for final-answer coverage and a trajectory-aware DSD-Score for tool-augmented search agents. Its Prerequisite Edge Score (PES) measures whether recorded tool trajectories follow annotated prerequisite edges. The dataset comprises instances spanning business domains, atomic claims spanning enumerations, numerical values, and factual descriptions (average per instance), and dependency edges (average per instance), with mean Width , Depth , and Members . Experiments show that state-of-the-art agent frameworks powered by frontier LLMs achieve low success rates on DSD, while trajectory-aware PES complements final-answer coverage with a prerequisite-edge execution signal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.