Model Merging Reshapes Learning Dynamics: A Functional Scaling Law Perspective
Abstract
We study model merging in power-law kernel regression, a tractable setting for understanding how checkpoint averaging reshapes learning dynamics. Specifically, we establish a functional scaling law (FSL) for fractional averaging in intrinsic time under general learning rate schedules, decomposing merged loss dynamics into signal learning and noise forgetting. Compared with the last iterate, merging increases signal error and can saturate its decay exponent. Its effect on noise depends on the injection time: merging retains more early noise but suppresses late noise, making its overall benefit depend on the learning rate schedule. The key mechanism is that averaging retains less-trained checkpoints, while late noise affects only the checkpoints following its injection. This characterization yields implications for merging design: tail and power averaging mitigate signal saturation, power averaging corresponds to a dynamic exponential moving average with a growing window, and sufficiently dense checkpoints preserve the noise-reduction rate of continuous averaging. Controlled SGD experiments in kernel regression empirically support the predicted data-scaling rates, optimal-window trends, and checkpoint-count criterion. Large language model (LLM) pretraining experiments show cross-schedule prediction without refitting and illustrate how the inferred signal–noise tradeoff explains observed merging gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.