SearchProof: Checkable Negative Evidence from Zero-Result Agent Search
Abstract
Search tools tell language-model agents what they found, but rarely what they actually covered. This asymmetry makes a zero-result call unusually difficult to reason about: an empty result can justify absence inside a searched directory while saying nothing about the rest of a repository. We introduce SearchProof, an inference-time interface that compiles a proposed negative claim into a claim-relative coverage obligation. Each search returns a compact receipt recording the relevant universe, executed scope, result set, coverage, and minimal completion set. The agent can pass this checked set directly to ordinary search before licensing a global negative conclusion. We evaluate this capability with ScopeBench: 72 controlled dependency-graph tasks and 144 query-matched tasks mined from 8 repositories spanning 5 language ecosystems and 2,625 source files. Every identifier appears once in each case type, and query-only balanced accuracy remains near chance. Across 4 agents and 1,872 analyzed model–task outcomes, agents recognize incomplete coverage reliably, yet exact controlled completion planning is only 25.0% from raw results and 100.0% with SearchProof. In interactive use, SearchProof raises grounded success from 94.8% to 100.0% in the two-model ablation grid. A preregistered full-corpus confirmation across DeepSeek and GLM-5.3 then moves 267/288 paired outcomes (92.7%) to 288/288 (100.0%): 21 wins, no losses. The direction remains positive at every nonzero session, query-group, and repository aggregation. Removing the executable completion field reduces DeepSeek to 95.1% (7–0 restored outcomes), isolating the operative mechanism. A source-backed deterministic executor reaches 144/144 tasks with exact completion and no follow-up language-model decision. Across 720 saved agent scopes, direct file execution agrees with the frozen backend on 100.0% of outcomes. SearchProof reads 62.0% fewer files and 60.6% fewer bytes than full-universe restart, reducing median scan latency by 54.3%. On the discriminating incomplete-absence stratum, the two-model gain is 79.2% to 100.0%. A third-agent stress replication moves from 41.7% to 85.4% (paired +43.8 points, 95% CI [25.0, 62.5]), with positive effects after holding out every repository. Median receipt construction plus checking is 4.53 ms at 4,096 components. These results make coverage a first-class evidence type and turn “not found” into a precise, executable observation for tool-using agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.