When Is Reuse Worth Replacing Exploration? Quality-Constrained Rollout Allocation for Stateful Agentic Reinforcement Learning
Abstract
Agentic reinforcement learning improves language models through multi-step environment interaction, but costly rollouts can yield little reward contrast, while restarting trajectories repeatedly discards reusable progress. Efficient training therefore requires deciding how much experience to acquire and when to generate, branch, or replay it. We propose QCAR, a quality-constrained rollout allocation framework for stateful agentic reinforcement learning. QCAR organizes observed task progress and explored transitions into a state graph, using initial Probe rollouts to estimate continuation success and identify valuable Fork boundaries. These estimates guide allocation across Fresh Generation, Fork, and Replay, with quality feedback limiting unproductive reuse. Balanced policy updates account for unequal sampling and shared prefixes, while observed outcomes recalibrate the predictor as the policy evolves. Across two models and four benchmarks, QCAR improves final accuracy over GRPO in all eight settings, averaging percentage points (% relative). QCAR yields up to % more reward-contrasting groups, with gains in all six model–environment comparisons, while fresh-token consumption decreases in five, by up to %. These results demonstrate the value of graph-informed, quality-constrained experience acquisition for stateful agentic reinforcement learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.