Endogenous Forgetting in RL Post-Training
Abstract
A central challenge in scaling reinforcement learning (RL) for LLM post-training is how models continue to learn new behaviors while preserving previously acquired ones. We study *endogenous forgetting*–where a model loses previously accessible behaviors due to non-stationarity that arises because the evolving model determines the effective data distribution–as a potential barrier to scaling on-policy RL. We first isolate endogenous forgetting via a controlled experiment: we collect successful traces from an RL run and show that supervised learning on these traces in temporal order leads to substantially more forgetting and lower solution coverage than training with a shuffled order. We then investigate successful-trajectory replay as a way to mitigate endogenous forgetting. Across multiple reasoning tasks and model families, replay reduces forgetting and prompt-level performance oscillations, achieves comparable or higher final pass@1, and substantially improves solution coverage relative to standard RL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.