Search-G1: Co-Evolving Representation-Based Rewards for Selective Grounded Search
Abstract
Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Yet terminal correctness assigns equal credit to direct answers from parametric knowledge, evidence-grounded answers, and answers recalled after redundant search. Process supervision and LLM judges provide richer feedback but can require high-cost annotation or inference; low-cost, non-contrastive policy confidence signals do not directly identify whether retrieved evidence changes an answer. We propose Search-G1, a reward framework that couples retrieval necessity with evidence reliance through two intervention-calibrated hidden-state readouts for selective grounded search. A prompt readout predicts closed-book sufficiency, whose complement proxies policy-relative retrieval necessity; an answer-commit readout predicts evidence reliance as sensitivity to evidence deletion conditional on the realized reasoning prefix. A correctness-first reward gates evidence-reliance credit by retrieval necessity, rewards sufficient direct answers, and penalizes repeated search. After calibration, scoring needs neither process annotations nor LLM-judge inference. As policy-relative targets, trajectory distributions, and representations may drift, Search-G1 periodically regenerates targets and refits its readouts. Across four QA benchmarks and two scales, Search-G1 achieves competitive task accuracy, with stronger counterfactual evidence dependence and fewer searches in the main comparisons, complemented by independent human audits of answer support.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.