Know Your Tokens: Structure-Aware Gradient Allocation for Agentic Search via Reinforcement Learning
Abstract
Reinforcement Learning (RL) from verifiable rewards is a powerful paradigm for training Large Language Model (LLM)-based search agents. Applying RL to agentic search models faces a core technical bottleneck: (a) standard Group Relative Policy Optimization (GRPO) causes gradient instability under parameter-efficient training (LoRA), and (b) existing fixes that subsample tokens by position (Stochastic GRPO, S-GRPO) restore convergence but discard gradient signal from late-position decisions that are equally critical. We show that position is not a well suited proxy for importance in multi-step search trajectories where critical decisions can occur at any position. We propose SA-GRPO (Structure-Aware S-GRPO), a general methodology for training LLM-based search agents via RL that addresses both problems: each model-generated token receives an inclusion probability in the policy gradient loss based on its functional role in the trajectory; search decisions and queries always receive gradient; answer tokens at high probability; reasoning is subsampled; and redundant re-searches are rarely included. Role assignment requires only deterministic parsing of tool-use tags, adding zero compute overhead. We evaluate on the HotpotQA and MuSiQue multi-hop Question Answering (QA) benchmarks using Mistral-7B-Instruct and Qwen2.5-7B-Instruct with QLoRA (NF4 quantization with LoRA adapters), comparing the QLoRA baseline against four RL-trained arms: GRPO, S-GRPO, SA-GRPO, entropy-weighted baseline under identical training conditions. Using a normalized answer-quality score (0-1) on Mistral/QLoRA, all subsampled arms improve over the baseline (0.620) while full GRPO barely moves (0.642). On HotpotQA, S-GRPO leads (0.703) with SA-GRPO close behind (0.684, 50% fully correct vs 40% baseline), under the same expected token budget (). On genuine 2-search trajectories under Qwen, where role-based selection provides maximum benefit, SA-GRPO scores 0.820 versus S-GRPO's 0.777. On the harder compositional questions of MuSiQue, SA-GRPO leads in both regimes: under the resource-constrained Mistral setting it is the only arm to improve over the baseline (0.416 vs 0.349, McNemar vs base), while GRPO falls below base and the entropy baseline collapses; under the stronger Qwen base all arms improve and SA-GRPO scores highest (0.590 vs 0.517) though within noise of GRPO (0.579). Beyond the method, we identify a key empirical finding: role-based selection requires multi-hop trajectory diversity to differentiate from position-based selection, a design guideline with implications for evaluation design in agentic RL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.