Feature Core Preservation: Protecting Task-Sensitive Parameter Subspaces for Model Merging
Abstract
Model merging has emerged as an efficient parameter-editing paradigm for augmenting pretrained models with fine-tuned downstream weights. A central challenge is that merging task updates can overwrite task-shared knowledge encoded in the pretrained model and degrade downstream performance. Core Subspace Preservation (CSP) mitigates this issue by protecting the dominant singular subspace of pretrained weights, along which the pretrained model is already locally optimal. However, CSP leaves directions highly sensitive to downstream task updates unprotected. In this work, we revisit subspace protection from a sensitivity perspective and show, both analytically and empirically, that directions with high downstream gradient energy have low perturbation tolerance and are therefore vulnerable to merging interference. To make such protection practical, we establish a connection between parameter gradients and layer input features, and propose Feature Core Preservation (FCP), which protects pretrained weights along dominant input-feature subspaces. FCP is a lightweight plug-and-play module that can be applied after arbitrary merging rules without requiring costly backward propagation. Extensive experiments on vision and language benchmarks, across different task scales and fine-tuning regimes, show that FCP consistently improves strong model merging baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.