Surgical Post-Training: Improving Reasoning While Mitigating Forgetting
Abstract
Injecting new reasoning knowledge into Large Language Models (LLMs) via post-training often induces catastrophic forgetting. Recent studies emphasize the importance of on-policy data in mitigating forgetting, but suggest that explicit KL regularization does not explain the observed retention benefits. We study the implicit regularization induced by reference-relative reward objectives: gradient analysis reveals adaptive attenuation of updates, while controlled comparisons on the same rectified data demonstrate improved capability retention. This motivates our Surgical Post-Training (SPoT), a post-training framework designed to optimize reasoning efficiently while preserving prior knowledge. SPoT consists of (1) a data rectification pipeline employing an Oracle to surgically correct erroneous steps via minimal edits, generating proximal on-policy data; and (2) a reward-based binary cross-entropy objective that learns from both corrected and erroneous responses. Empirically, with only 4k rectified math pairs, SPoT improves Qwen3-8B's accuracy by 6.3 percentage points on average across 9 in-domain and out-of-domain tasks, requiring merely 16 minutes of model training on 8× H800 GPUs. We further validate SPoT through Connect4 post-training, where it improves both game and mathematical reasoning while retaining instruction-following capabilities. Moreover, SPoT provides a superior initialization for subsequent reinforcement learning, enabling further gains in reasoning performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.