acceptodds
Under review as a conference paper at ICLR 2027

Never Overshoot: Expert-Anchored Gradient Descent for Model Merging

Abstract

Model merging combines task-specific experts fine-tuned from a common pretrained model into a single model, which often underperforms the experts on their respective tasks. We study refinement using a few labeled examples per task, already available for selecting or fitting the merge. We introduce expert-anchored gradient descent (XGD), which uses the original expert checkpoints to guide updates from these small buffers. XGD scales each coordinate's learning rate by its distance to the task expert and permits only movements toward that expert, without passing it. Replay gradients select locally favorable movements, and expert offsets are recomputed as refinement proceeds. On GLUE-8, CLIP-8 and CLIP-20 with 8 or 16 labels per task, both XGD and AdamW improve every evaluated merge in mean task performance, reusing the examples already used to select or fit it. Thus, few-shot refinement offers gains without additional labels. Across five buffer draws, XGD has the best mean among the compared refinements in five of six benchmark–budget settings. Starting directly from the pretrained model, XGD also outperforms many merge baselines, combining expert information without a separate merging algorithm. From the same merges at eight labels per task, XGD retains higher mean ImageNet accuracy than AdamW in every comparison and lower mean WikiText-2 masked-LM loss in all but one.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.