Search Environment Agent for Efficient Deep Search Reinforcement Learning
Abstract
Training deep search agents with reinforcement learning (RL) on the real web is costly, motivating cheaper simulated environments. Yet we observe **reward–transfer decoupling**: training reward in the local environment improves without corresponding real-web gains. Even factually correct feedback can omit evidence obtainable by effective queries or over-assist underspecified ones, distorting which search behaviors RL reinforces. We introduce **Search Environment Agent (SEA)**, a memory-augmented environment manager that calibrates action-dependent access to evidence. Guided by **Oracle Memory**, its internal record of real-tool experience, SEA selects and reranks locally supported evidence and reserves real execution for predicted decision-relevant discrepancies that cannot be repaired locally. Paired local and real feedback supervise response selection and routing, with counterfactual policy continuations informing routing labels. Verified real observations update the memory and webpage index. Evaluated exclusively with real-web tools without SEA or Oracle Memory, SEA-RL-35B surpasses reported specialized-agent baselines across six benchmarks. In controlled 4B comparisons, SEA outperforms Cached-Web RL on all four benchmarks and Real-Web RL on three, using approximately 4.8% of the latter's measured training API expenditure. Diagnostics further show improved evidence recall alongside reduced over-assistance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.