acceptodds
Under review as a conference paper at ICLR 2027

Learning Not to Interfere: Adapting Experts for Model Merging

Abstract

Model merging combines experts fine-tuned from a shared pretrained initialization into a single multitask model, but cross-task interference often degrades their specialized capabilities. We argue that this interference stems from how experts are fine-tuned. Standard fine-tuning optimizes each expert on its own task distribution without constraining how its parameter update affects predictions on other tasks. These parameter updates, known as task-vectors, represent changes from the shared initialization and can disrupt other experts when combined. We propose Resolving Interference (RI), a lightweight framework that independently adapts experts before merging using unlabeled task data. RI distills the original expert's predictions on its own task to preserve specialization, while matching the shared pretrained model's predictions on other tasks. The pretrained model provides a reference for behavior before specialization, allowing RI to restrict each task vector's influence beyond its own task distribution. RI complements existing merging methods, including those using test-time adaptation, and improves performance across vision and language benchmarks. For language-model merging, RI surpasses the strongest evaluated baselines by up to 2.9%. Additionally, quantitative consistency evaluations and qualitative generations confirm RI's ability to better preserve the original experts' behavior across tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.