acceptodds
Under review as a conference paper at ICLR 2027

Endogenous Forgetting in RL Post-Training

Abstract

A central challenge in scaling reinforcement learning (RL) for LLM post-training is how models continue to learn new behaviors while preserving previously acquired ones. We study *endogenous forgetting*–where a model loses previously accessible behaviors due to non-stationarity that arises because the evolving model determines the effective data distribution–as a potential barrier to scaling on-policy RL. We first isolate endogenous forgetting via a controlled experiment: we collect successful traces from an RL run and show that supervised learning on these traces in temporal order leads to substantially more forgetting and lower solution coverage than training with a shuffled order. We then investigate successful-trajectory replay as a way to mitigate endogenous forgetting. Across multiple reasoning tasks and model families, replay reduces forgetting and prompt-level performance oscillations, achieves comparable or higher final pass@1, and substantially improves solution coverage relative to standard RL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.