Search-Agent Scaling Gains Survive Answer-Witness Deletion
Abstract
Search agents are evaluated by answer accuracy and by whether retrieved evidence supports their answers, but neither reveals how an answer responds when that evidence is removed. We introduce counterfactual evidence accounting: at every retrieval call, passages matched as containing an accepted-answer alias are deleted and replaced by lower-ranked candidates, and the complete search policy is rerun. Paired rollouts partition ordinary accuracy exactly into correct answers that become incorrect after this deletion (D-failing) and those that remain correct (D-surviving). On 623 matched questions, the released 14B Search-R1 checkpoint is 10.3 percentage points more accurate than the 3B checkpoint, and 9.5 of those points are D-surviving; the analysis detects no increase in D-failing correctness. Five released Qwen2.5 search-agent recipes, prompted Qwen checkpoints without search-agent reinforcement learning, and an SSRL Llama comparison show the same concentration, which persists under residual-witness, parsing, and exposure sensitivity analyses. With matched weights, 7.9 of the 9.5 points overlap with closed-book correctness. Deleting passages that match the agent's own answer yields a label-free diagnostic, self-ablation: with one additional rollout, it lowers D-surviving correctness among released answers from 8.1% to 2.4% at 81% coverage, at a cost of 8.3 points of total correct yield. Search-agent evaluations should report how accuracy responds to targeted changes in retrieved evidence, alongside accuracy and support.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.