Mind the Domain Gap: Post-Training Robustness Is Geometrically Selective
Abstract
Post-training can acquire new capabilities at the cost of forgetting existing ones, yet how this forgetting depends on the training objective remains poorly understood. We study rebound—the recovery of pretraining behavior under opposing updates—as a probe of post-training robustness. Starting from the same Qwen3-4B-Base, we train matched mathematical capabilities with SFT and GRPO under nearly identical update-token budgets, and then apply controlled interference ranging from same-domain math to cross-domain code and far-domain summarization. We find that RL's apparent advantage in retention is strongly domain-dependent. GRPO retains more under cross- and far-domain interference (84.3% vs. 78.0%; 95.0% vs. 72.6%), but offers no advantage under same-domain interference (49.7% vs. 50.0%). This reversal is explained by the geometry of the learned updates: GRPO concentrates its change in a low-dimensional subspace, whereas SFT spreads it more broadly. Interventions on the update spectrum establish causality: GRPO's dominant directions are both sufficient and necessary for its retention advantage, while SFT contains a harmful low-energy tail. These results show that retention is not an intrinsic property of the post-training objective. Rather, forgetting is governed by the alignment between subsequent interference and the capability-bearing subspace induced by post-training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.