CFA: Mitigating Forgetting in RL Post-Training with Content-Free Anchors
Abstract
RL post-training is deployed under the assumption that it forgets little of what it was not trained on. We show this is incomplete: across Qwen3-VL at 2B, 4B, and 8B scale, single-domain RL post-training silently collapses held-out capability, with worst-decile drops up to −67% hidden behind suite-wide averages that stay flat only because unrelated benchmarks improve on the same run. We introduce the worst-k curve to expose this tail. Every existing defense, parameter-importance penalties, replay, KL toward a reference, assumes the anchor set used to constrain the model must represent the real capabilities being protected. Protecting a handful of anticipated domains this way is expensive; protecting general capability this way is infeasible, since it requires anchoring on essentially the model’s entire pretraining distribution. We find this assumption is not load-bearing: anchors drawn from pure random noise, with no connection to any capability, match and at modest budgets outperform a curated real coreset spanning five capability dimensions. We instantiate this as CFA, a Fisher-weighted Maximum Mean Discrepancy penalty between current and reference next-token distributions, which we derive as a Gauss-Newton generalization of Elastic Weight Consolidation whose finite-sample guarantees hold regardless of anchor content. Across all nine domain-scale combinations tested, CFA holds worst-decile forgetting to −0.3% to −9.1%, against −8.1% to −67.4% for the best baseline, while matching or exceeding in-domain gain, using no real, licensed, or curated data at all.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.