One Answer to Grade Them All? Reference-Based Evaluation and Its Optimization Effects in Search Agents
Abstract
Search agents have become increasingly capable of tackling complex questions, with reliable judgments of answer correctness providing the basis for assessing their capabilities and guiding further learning. Search-agent evaluation and training commonly rely on reference answers, although recent work incorporates verified alternative answers into training rewards. However, the implications of incomplete reference coverage for evaluation and supervision remain insufficiently understood in more challenging web-search tasks. Questions combining multiple constraints may still admit valid solutions beyond the references, causing erroneous rejection and potentially misleading supervision. To examine this problem, this study uses evidence-based verification to quantify grading discrepancies and analyzes their sources through question constraints and supporting evidence. Policy updates under original and corrected rewards are then compared to examine how grading errors affect supervision of answers and search actions. Extensive experiments across five search benchmarks show that reference-based evaluation can underestimate performance by rejecting valid answers, with verified revisions increasing the acceptance rate by 2.12 pp. Source-based attribution identifies reference omissions as a distinct error category arising from one-to-many relations or unspecified selection rules. When used as rewards, these misjudgments can lower the likelihood of valid answers and useful search actions, depending on response content, group composition, and search context. These findings motivate checking for competing valid solutions when constructing questions and references for evaluation and learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.