acceptodds
Under review as a conference paper at ICLR 2027

RL Does Not Always Forget Less: Understanding Retention through Target Learning in LLM Post-Training

Abstract

Reinforcement learning (RL) often preserves prior capabilities better than supervised fine-tuning (SFT) during post-training, but we show that this retention advantage is not universal and can reverse. We study this variation by connecting target learning to non-target retention through gradient analysis and controlled experiments across diverse models, target tasks, and learning rate settings. We first decompose the SFT gradient into a capability component that increases correct-response probability and a supervision component that fits the distribution over provided correct responses. Under binary rewards, the expected on-policy RL gradient is proportional only to the capability component. This distinction reveals how the initial model affects target learning differently under RL and SFT. For RL, low initial correct-response probabilities can limit reward variation and restrict target learning at small learning rates. For SFT, compatible supervision can enable comparable target gains at smaller learning rates and, in some settings, better retention than RL. We further show that the learning rate regime is closely tied to non-target retention: at small learning rates, degradation is typically localized and cross-entropy changes are largely explained by first-order gradient update interactions, whereas at larger learning rates, nonlinear contributions can become substantial and degradation can extend across more non-target tasks. Together, these results explain when RL's retention advantage holds or breaks down and identify the roles of the initial model, target learning efficiency, and learning rate regime in the relative retention of RL and SFT.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.