Learning Not to Interfere: Adapting Experts for Model Merging
Abstract
Model merging combines experts fine-tuned from a shared pretrained initialization into a single multitask model, but cross-task interference often degrades their specialized capabilities. We argue that this interference stems from how experts are fine-tuned. Standard fine-tuning optimizes each expert on its own task distribution without constraining how its parameter update affects predictions on other tasks. These parameter updates, known as task-vectors, represent changes from the shared initialization and can disrupt other experts when combined. We propose Resolving Interference (RI), a lightweight framework that independently adapts experts before merging using unlabeled task data. RI distills the original expert's predictions on its own task to preserve specialization, while matching the shared pretrained model's predictions on other tasks. The pretrained model provides a reference for behavior before specialization, allowing RI to restrict each task vector's influence beyond its own task distribution. RI complements existing merging methods, including those using test-time adaptation, and improves performance across vision and language benchmarks. For language-model merging, RI surpasses the strongest evaluated baselines by up to 2.9%. Additionally, quantitative consistency evaluations and qualitative generations confirm RI's ability to better preserve the original experts' behavior across tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.