acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking and Training Reward Models for Cross-Source Evidence Grounding in Search Agents

Abstract

Existing approaches to training and evaluating search agents focus heavily on final-answer quality, largely neglecting how generated answers map back to retrieved evidence. Recent methods incorporate evidence citation tracking to enforce grounding, but citation support alone does not establish source independence or whether an answer appropriately handles conflicting evidence. In this paper, we argue that credibility in open-web retrieval depends on principled cross-source evidence grounding that accounts for source independence and systematically attributes and adjudicates conflicting evidence, while hierarchical judgments of these grounding decisions provide a unified signal for training and evaluation. To this end, we propose an evaluation framework, **Cross-Source Evidence Grounding (CSEG)**, that employs a fine-grained, six-layer decision tree over the complete search trajectory. We introduce **CSEG-Bench** and trainable reward models to provide targeted supervision following this framework. Our method offers the key advantage of eliminating the severe data skew of natural trajectories by providing targeted coverage over complex, deep-layer decision paths, such as cross-source conflict resolution and nuanced attribution, that standard rollouts fail to capture. Specifically, (1) we define the hierarchy at label, node, and path levels; (2) we construct 1,160 verified trajectories over nine canonical paths by rewriting answers on real retrieved document pools, with independent frontier judges reproducing the entire gold path and a human expert deciding inclusion; and (3) we train **CSEG-RM**, which reaches 83.9% exact accuracy on our benchmark while lying on the performance–token efficiency Pareto frontier. Experiments that use this reward alone to supervise a frozen-retrieval answer policy raise the independently judged mean label from 2.83 to 4.19 and show that learning the hierarchy generalizes across layers at a fraction of frontier-judge cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.