Incremental Spectral Merging with Pre-trained Models
Abstract
Despite much recent progress, continual learning remains unsolved. Every update to a model's weights risks erasing what the model already knows. Model merging is a promising way to address this challenge. Treating each task update as a separate model, merging combines multiple models into one without retraining and without storing data or extra models. Even simple operators such as averaging or maximum-magnitude selection work well in practice. However, there is little theory explaining why the merged model performs well in continual learning. In this paper, we offer a new perspective on model merging for continual learning. Instead of merging all parameters of the low-rank modules used for fine-tuning, we merge only the dominant singular directions of each update, scaled by a damping factor. Our theoretical analysis bounds the loss of the merged model by two terms: the change of the classification head and the weight step between consecutive modules. Spectral merging sets the second term through its damping factor, and realigning the classifier after the merge keeps the first small. Our experiments show that keeping the dominant low-rank information is key to mitigating forgetting. Across four class-incremental benchmarks, our method ranks first in average final accuracy and in robustness to input corruptions, ahead of both prior continual learning methods and published merging rules.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.