acceptodds
Under review as a conference paper at ICLR 2027

Selecting What to Learn and Correcting Policy Gradients for Search Agents

Abstract

Search agents extend language models to information-seeking tasks that require repeated retrieval and reasoning across sources. Learning from rewards based on final-answer correctness remains difficult. Questions that are too easy or too hard provide little reward contrast, while sampled trajectories produce noisy gradients. We introduce an adaptive training framework that couples question selection with gradient correction. A frozen generator constructs questions from retrieved evidence, and independent probes estimate the current agent's success rate on each question. The framework then routes questions to supervised fine-tuning, reinforcement learning, or later reassessment. For reinforcement learning, we propose Gradient-Corrected Group Relative Policy Optimization (GC-GRPO) to reduce sampling noise in policy gradients. GC-GRPO constructs zero-mean reference gradients from current trajectories and combines them using weights learned from earlier batches. Our analysis shows when this correction reduces mean squared estimation error while preserving the expected GRPO gradient at the sampling policy. Across five search benchmarks, the trained 4B search agent achieves 85.44% accuracy on GAIA, 48.66% on BrowseComp, 47.75% on BrowseComp-ZH, 47.75% on SEAL-0, and 82.00% on XBench-DS.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.