acceptodds
Under review as a conference paper at ICLR 2027

Measure the Merge in the Currency You Deploy: Calibrated Path Decisions on a Budget

Abstract

Merging two fine-tuned models by weight averaging can quietly destroy the task: averaging two capable, officially released Qwen2.5 specialists can leave a model that solves nothing, while another pair from the same family survives — and nothing readable from the endpoints, from task labels to weight distance to training logs, says which is which. Part of that cost is exact rather than empirical: the conventional 50/50 average forfeits half the endpoint gap on any asymmetric pair, because net gain over the better endpoint decomposes pointwise as dip depth minus an endpoint-asymmetry wedge. The rest must be measured along the interpolation path, but path evaluations cost GPU time: how many to buy, and what to read them in? How many: a midpoint accept/reject decision is a three-evaluation problem — three points plus one calibrated safety margin, which transfers unchanged to three external populations and is measured wider on the fourth; a denser grid raises coverage on paper but admits no more pairs at the same calibrated risk. For finding a coefficient, density buys real structure at a minimax-optimal k⁻² rate that flattens past about six queries. In what currency: at matched wall-clock budget, rules that read the proxy-loss curve repeatedly committed coefficients below the better endpoint — the worst landing −17.0 points below it — while rules that measured the task metric directly did so far less often; where paths are flat, nothing separates the two currencies. The coefficient, not the recipe, is the operative variable: TIES and DARE collapse at full strength and recover at the measured coefficient where the path allows it. Read the path's depth once, then measure the merge in the currency you deploy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.