Forgotten or Inaccessible? Restoration Savings Reveal What Language Models Retain After Fine-Tuning
Abstract
Catastrophic forgetting is commonly defined as a substantial drop in performance on previously learned tasks after training on new tasks or data. However, this behavioral failure does not establish that the underlying knowledge was erased. We test this distinction with restoration savings, which asks whether, given the same small amount of retraining, a forgotten ability is easier to recover than a comparable ability that was never learned. We teach a model a random subset of facts it did not know, induce forgetting by fine-tuning it on an unrelated task, and then retrain learned and never-learned facts under this same budget. Previously learned facts are recovered far more easily, even when they are not themselves retrained and when they are asked in new wordings. Together, these findings show that behavioral access can be lost while a recoverable trace of prior learning remains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.