Transfer: Coefficient Transfer for Efficient Model Merging
Abstract
Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as model size and number increase, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly similar performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{\alphaTransfer}: search for optimal coefficients on a small proxy model, then directly transfer them to larger target models in the same family. We verify Transfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6 speedup and 70% memory reduction on vision transformers, and a 20 speedup and 85% memory reduction on large language models, while maintaining comparable performance. Our findings establish Transfer as an efficient and generalizable approach to scaling model merging.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.