acceptodds
Under review as a conference paper at ICLR 2027

What Is a Merge Worth? Data and Training Budgets in Model Merging

Abstract

Model merging can produce a checkpoint for immediate deployment or an initialization for further adaptation. These uses raise different questions: a merged model's deployment advantage may change when the pretrained base is also allowed to train, and single-budget evaluations cannot separate the value of initialization from access to more data or optimization. We compare fixed TSV and pretrained CLIP initializations across eight classification tasks, using matched adaptation and paired seeds to measure changes in their performance gap along data and update axes that share a common endpoint. On ViT-B/32, increasing the fitting pool reduces the gap more than extending training on the full pool; on ViT-B/16, the data contrast points in the same direction but remains uncertain. TSV retains an advantage at the shared endpoint, while supervised and magnitude controls show that the result is not limited to distillation or explained by update norm alone. The findings suggest that practitioners compare merged and pretrained starts across adaptation budgets and interpret merge value within the evaluated access regime, rather than as a universal trade-off or end-to-end saving.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.