Specialist Drift: Continual RL for Capability-Preserving Foundation Model Post-Training
Abstract
We ask whether reinforcement-learning (RL) post-training of foundation models is better understood as single-task optimization or as an unrecognized continual RL (CRL) problem. Existing pipelines (RLHF, RLVR, agentic RL) report target-capability gains without characterizing which prior capabilities silently decay, or by how much. We present a parameter-matched causal decomposition of this specialist drift, isolating distribution narrowing (61.1%), reward gradient asymmetry (17.8%), token-entropy collapse (13.9%), and KL anchor decay (7.2%) as distinct contributors to capability erosion across four post-training stages and three open-weight backbones. Building on this decomposition, we propose Subspace-Constrained Policy Optimization (SCPO), which retains 92.1% of base-model capability tomography while matching 99.1% of the unconstrained PPO target reward on competition math, a 6.4 improvement in retention per unit of target gain over global KL anchoring. We further introduce two preservation metrics, CPS and DRI, along with a fitted scaling law (S.D.S.L.) relating subspace-anchor strength to capability retention as the training horizon grows (, ), showing that behavioral subspace structure, rather than parameter count or KL budget, governs the retention frontier in long-horizon RL post-training. Extended experiments include a new strongest baseline (Heterogeneous Model Averaging), modern continual-RL preservation baselines (CPO, RPO), divergence-regularized baselines (DRPO/DPPO), SFT/replay-SFT controls, out-of-sample validation of the scaling law, and a Fisher-information verification protocol confirming the theoretical predictions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.