Two Checkpoints Are Enough: Data-Free Post-Hoc Repair of Catastrophic Forgetting
Abstract
Fine-tuning a language model for a target task degrades capabilities that the training data never explicitly threatened, yet it also improves other held-out skills as a useful side effect. Low-rank fine-tuning methods limit the forgetting but forfeit these incidental gains. Furthermore, recovering the damaged capabilities usually requires the data the model was originally trained on, which practitioners who fine-tune open-weight models rarely have. We study post-hoc repair of this catastrophic forgetting using only the pre-fine-tuned checkpoint and its fine-tuned descendant . The goal is to recover the damaged capabilities while preserving the target-task gains and incidental held-out improvements that reversion toward the base discards. We show that a single fine-tuning update separates in singular-value space into a small set of task-aligned spikes and a broad bulk that matches the Marchenko-Pastur prediction for an IID noise matrix. DG-Hard applies the Gavish-Donoho optimal hard singular-value threshold to each weight-delta matrix, keeps the spikes, and reverts the bulk. The cutoff is closed-form, so the repair needs no further training, gradients, data, or hyperparameters. To evaluate such methods, we also introduce a partition-conditional metric that separately tracks healing, preservation, non-damage, and target-task retention. Across nine held-out benchmarks, DG-Hard attains both the highest combined score and the highest mean held-out accuracy among non-spectral training-time and post-hoc baselines. We show that the retained signal carries the incidental held-out gains, and a model rebuilt from the discarded bulk alone keeps only a third of the incidental held-out gains. Alternative truncation rules such as fixed-rank and fixed-energy truncation reach or exceed DG-Hard only when their retained fraction is tuned against old-task performance, while DG-Hard's closed-form cutoff needs no tuning. The same repair also recovers most of the safety alignment eroded by benign fine-tuning, without alignment data. The code is available at https://anonymous.4open.science/r/dghard-repair-2D84/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.