Selecting What to Learn and Correcting Policy Gradients for Search Agents
Abstract
Search agents extend language models to information-seeking tasks that require repeated retrieval and reasoning across sources. Learning from rewards based on final-answer correctness remains difficult. Questions that are too easy or too hard provide little reward contrast, while sampled trajectories produce noisy gradients. We introduce an adaptive training framework that couples question selection with gradient correction. A frozen generator constructs questions from retrieved evidence, and independent probes estimate the current agent's success rate on each question. The framework then routes questions to supervised fine-tuning, reinforcement learning, or later reassessment. For reinforcement learning, we propose Gradient-Corrected Group Relative Policy Optimization (GC-GRPO) to reduce sampling noise in policy gradients. GC-GRPO constructs zero-mean reference gradients from current trajectories and combines them using weights learned from earlier batches. Our analysis shows when this correction reduces mean squared estimation error while preserving the expected GRPO gradient at the sampling policy. Across five search benchmarks, the trained 4B search agent achieves 85.44% accuracy on GAIA, 48.66% on BrowseComp, 47.75% on BrowseComp-ZH, 47.75% on SEAL-0, and 82.00% on XBench-DS.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.