Watch Where You Linearize: Predicting and Correcting Model-Merging Interference
Abstract
Model merging composes separately fine-tuned models by adding their task vectors to a shared backbone. In standard merging pipelines, the resulting functional interference is known only after candidate merges are built and evaluated. We show that a single directional derivative, the receiving model's output response to another task's vector, forecasts the interference before any candidate is evaluated, predicting logit drift, label flips, per-example damage, and per-task accuracy of the merge. The forecast works only in the receiver's local geometry: read at the shared pretrained initialization, the same derivative is nearly useless, because fine-tuning reorients the cross-task responses at the very beginning of training. Freezing the receiver-local responses yields a surrogate whose outputs are affine in the merge coefficients, and with it LR-Merge, a label-free correction that selects coefficients through a sequence of convex teacher-matching problems on task-specific unlabeled probes, never backpropagating through a candidate merged network. On vision backbones and language models, LR-Merge improves every strong base merger we test. Its advantage concentrates where merging is data-scarce or repeated. At low probe budgets it surpasses strong coefficient-optimization baselines given several times more data and compute, and the frozen responses are built once and reused, so each further coefficient setting or task subset costs seconds rather than a rebuild.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.