acceptodds
Under review as a conference paper at ICLR 2027

When Is Reuse Worth Replacing Exploration? Quality-Constrained Rollout Allocation for Stateful Agentic Reinforcement Learning

Abstract

Agentic reinforcement learning improves language models through multi-step environment interaction, but costly rollouts can yield little reward contrast, while restarting trajectories repeatedly discards reusable progress. Efficient training therefore requires deciding how much experience to acquire and when to generate, branch, or replay it. We propose QCAR, a quality-constrained rollout allocation framework for stateful agentic reinforcement learning. QCAR organizes observed task progress and explored transitions into a state graph, using initial Probe rollouts to estimate continuation success and identify valuable Fork boundaries. These estimates guide allocation across Fresh Generation, Fork, and Replay, with quality feedback limiting unproductive reuse. Balanced policy updates account for unequal sampling and shared prefixes, while observed outcomes recalibrate the predictor as the policy evolves. Across two models and four benchmarks, QCAR improves final accuracy over GRPO in all eight settings, averaging percentage points (% relative). QCAR yields up to % more reward-contrasting groups, with gains in all six model–environment comparisons, while fresh-token consumption decreases in five, by up to %. These results demonstrate the value of graph-informed, quality-constrained experience acquisition for stateful agentic reinforcement learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.