PALM-Merge: Preserving Task Residuals for Data-Free Continual Model Merging
Abstract
Data-free continual model merging can continually extend deployed foundation models without retaining task data or retraining, but it must integrate new checkpoints while preserving earlier capabilities. Existing methods fold successive task deltas into shared weights and control interference by projection or filtering. However, forcing heterogeneous responses into one map creates an approximation floor and can overwrite directions required by earlier tasks. To address this limitation, we propose PALM-Merge (Preserve-and-Activate Low-rank Merging), which stores each task residual as an independent low-rank expert and activates experts sparsely from inputs, separating response preservation from access and avoiding cross-task parameter overwriting. Specifically, truncated SVD converts each task update into a frozen low-rank expert whose compression error is controlled by the discarded singular energy, while routing each activation to the most relevant low-rank expert at inference without access to any training data. Theoretically, we show that a historical prediction can change only when a newly added expert changes the routing decision. Forgetting can then occur only if the resulting perturbation is large enough to overturn the original classification margin. This yields a cumulative accuracy-loss bound for task \(c\) of \(\min{1,\sum_t=c+1^Tp_t,c}\). Across nine FusionBench CLIP-ViT settings, achieves \(89.8%\) average accuracy. On a controlled eight-task stream, it outperforms NUFILT by \(4.85%\) and reduces positive final forgetting by \(65.9%\), while reaching \(82.75%\) on Flan-T5 GLUE-6. These gains require \(2.10\times\) storage but only \(1.24\times\) active-tensor cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.