Fine-Tune to Merge: Spectral Core Preservation for Continual Learning
Abstract
Continual learning aims to enable large pretrained models to acquire new knowledge while retaining previously learned capabilities, yet remains challenged by catastrophic forgetting during sequential adaptation. Existing model-merging-based approaches mitigate forgetting by integrating model updates across tasks, but most still rely on standard fine-tuning and largely overlook the resulting perturbations to the pretrained core subspace. We find that old-task loss is more sensitive to parameter perturbations along pretrained core directions than along residual directions, making new-task updates along these directions more likely to interfere with previously learned knowledge. To address this issue, we propose Spectral Core-Preserving Optimization (SCPO), a fine-tuning method that protects the pretrained core subspace when learning new tasks and produces task updates better suited to model merging-based continual learning. SCPO defines the protection scope using the dominant singular subspaces of the pretrained weights and adaptively modulates the protection strength according to their overlap with the corresponding subspaces of the previous merged model, thereby selectively attenuating gradient components along protected directions. The filtered gradients are further optimized within an online-updated low-dimensional subspace, and the resulting task vectors are subsequently integrated through model merging. Extensive experiments on vision models and large language models demonstrate the effectiveness of SCPO. Across five representative model merging methods, SCPO consistently improves over standard fine-tuning, yielding an average relative accuracy gain of 5.44% and a maximum gain of 8.63%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.