Nostalgia: Continual Learning through Constrained Optimisation in Small Models
Abstract
Pre-trained transformer models are typically adapted to downstream tasks via supervised fine-tuning, but sequential adaptation often leads to **catastrophic forgetting**, where performance on earlier tasks degrades. Existing mitigation strategies commonly rely on data replay, architectural modifications, or heuristic regularization, limiting their scalability. We introduce *Nostalgia*, a data-free, optimization-based method for continual supervised fine-tuning that mitigates forgetting by constraining parameter updates to directions that minimally interfere with previously learned tasks. Our approach is grounded in a local quadratic approximation of task losses and operationalizes this principle through a tractable low-rank approximation of the average Hessian, enabling efficient projection-based updates compatible with modern parameter-efficient fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.