acceptodds
Under review as a conference paper at ICLR 2027

River-Hill Budgets: Understanding the Effects of Annealing in Model Growth

Abstract

Model growth initializes a larger model from a smaller source model that has already been trained, but saves compute only if the grown model properly reuses that source’s training. Work on growth has focused on the growth operator, which is in charge of preserving the source model’s function across the transfer while ensuring the newly added parameters are properly made use of. However, when param eters are transferred and the learning-rate schedule is reset, the sudden learning-rate increase creates a training offset on top of the parameter transfer. We separate the impact of the growth operator from the learning-rate increase and find that the newly expanded parameter space does not necessarily help recover from this schedule offset. Using the river-valley view of the loss landscape, we explain and quantify this offset. For a standard BERT configuration, the river-valley view predicts that 45% of progress made during the decay to 10% peak learning rate is lost upon reset, and up to 50% for a decay to zero. The same model trained with 45% fewer anneal steps, making only reset-surviving progress, closely matches the annealed run’s loss trajectory. These results show that the source model’s learning-rate schedule severely limits how much training survives model growth.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.