AdaLMM: Adaptive Layer-wise Model Merging for Continual Learning
Abstract
Replay-based continual learning mitigates catastrophic forgetting by retaining historical examples, but its gains can diminish as the replay memory grows. Parameter-space model merging offers a complementary solution, yet fixed global interpolation coefficients provide limited flexibility in balancing stability and plasticity across the network. To address this limitation, we propose Adaptive Layer-wise Model Merging (AdaLMM), a task-boundary merging framework for replay-based continual learning. AdaLMM first aligns the historical and updated models through activation-based Optimal Transport matching and architecture-compatible symmetry transformations, and then learns layer-wise or block-wise interpolation coefficients directly from replay data. This allows different parts of the network to retain different proportions of the newly learned update. Experiments on Seq-CIFAR-100 and Seq-Tiny-ImageNet with ResNet-18 and Vision Transformer backbones show that AdaLMM improves the balance between retaining previous knowledge and learning new tasks, leading to higher overall accuracy than standard replay across different memory budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.