Continual Learning in Agentic RL via Success-Constrained Policy Gradient Projection
Abstract
Agentic reinforcement learning can forget previously successful behaviors even within a single training run on a fixed task distribution. We quantify this trajectory-level forgetting, revealing a continual learning problem arising without an explicit sequence of tasks. At the trajectory level, maximum-return oracle trajectories have nonnegative advantage relative to the current policy's expected return, regardless of which policy generated them. This provides a policy-improvement motivation for selecting historical successes as protection targets. We introduce Success-Constrained Policy Gradient Projection, which makes the minimum Euclidean correction to the main loss gradient needed to satisfy an aggregate success-replay constraint. Our analysis characterizes the tradeoff between protecting historical successes and optimizing the current objective: for a gradient-descent step, the correction removes a first-order increase in aggregate replay loss, with a cost to the main objective determined by gradient alignment. Averaged over three training seeds on WebShop, the method improves success by 6.2 and 2.1 percentage points for GRPO and GiGPO, respectively, without changing their advantage estimators. GRPO with the correction reaches 71.4% mean success, surpassing the GiGPO baseline at 70.2%. In a supplementary ALFWorld comparison using one training seed without policy collapse, the method improves final success and reduces the decline from validation-selected saved checkpoints with both algorithms. These results demonstrate that retaining historical successes can yield gains comparable to finer-grained advantage estimation, and establish the practical value of success-based protection for continual learning in agentic reinforcement learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.