acceptodds
Under review as a conference paper at ICLR 2027

Scale First, Then Shape: The Predictive Geometry of Merging Multimodal Experts

Abstract

Model merging combines fine-tuned experts into a single model without training data. For multimodal large language models (MLLMs), where each expert adds one capability such as OCR or grounding to a shared base, it avoids assembling instruction data for every new capability. The merging methods proposed for this setting have grown steadily more elaborate, from sparsification and spectral editing to per-layer optimization, yet each is judged only by its final benchmark score. As a result, a practitioner must evaluate every candidate merge, at over an hour per evaluation, and no existing analysis says in advance which method will work on a given model, or why. In this paper, we show that for MLLM capability merging the answer can be read from the merged weights before any evaluation. (i) We introduce a two-axis geometric account of a merged displacement: an effective scale, how far the merge travels in units of the naive task-vector sum, and a signed depth slope, whether that travel tilts toward deep or shallow layers. Keeping only each method's component along the naive sum and evaluating a plain task-arithmetic curve at its measured scale predicts the benchmark average of three published methods (OptMerge, WUDI and Iso-C) on Qwen2-VL-7B and InternVL2.5-1B within one point in five of the six cases; the depth slope, calibrated on synthetic probes at fixed scale, accounts for the sixth. Re-merging the collapsed Iso-C model from that component alone lifts it from 18 to 47 points, within a point of where the curve placed it. (ii) We show that per-layer optimization reduces to this geometry: 91% of its displacement lies along the naive sum, and a closed-form profile matches the projected optimizer within 0.13 points on both MLLMs at one twentieth of the merging cost. (iii) We find that the component the projection removes adds 0.3 to 1.4 points for the optimizers on capability experts but carries 21 to 47 points when experts have disjoint label spaces, which marks the boundary of the account. Merging multimodal experts is a prediction problem rather than a search problem, and scale is the first quantity to get right.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.