acceptodds
Under review as a conference paper at ICLR 2027

Spectral Components Require Distinct Coefficients in Model Merging

Abstract

Model merging combines task-specific experts into a single multi-task model without joint retraining. Cross-task orthogonalization eliminates overlap between the retained task subspaces in weight space, yet the merged model's predictions still differ from those of the task experts. We show that reweighting retained spectral components reduces this expert discrepancy while keeping their directions fixed. We find that spectral components from the same task and layer affect expert discrepancy differently and therefore require distinct coefficients. Accordingly, we jointly learn component-wise spectral coefficients in a fixed, task-structured orthogonal basis by matching expert predictions on unlabeled calibration inputs. Across three CLIP backbones and benchmarks of 8, 14, and 20 tasks, our method achieves the highest accuracy among the evaluated methods in all nine settings. On ViT-B/32, it improves normalized accuracy over the strongest baseline by 3.2, 4.0, and 6.3 percentage points with 8, 14, and 20 tasks, respectively. Experiments on MergeBench further show improvements over the evaluated merging baselines for both base and instruction-tuned Llama-3.2-3B models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.