acceptodds
Under review as a conference paper at ICLR 2027

Double descent in learning rate in the feature learning regime

Abstract

Double descent is the non-monotonic behavior of test error as a function of a single hyperparameter. It is well documented for model width, dataset size and training time. We show that it also happens along a different axis: the learning rate. We give a mechanistic account of how this happens and under what conditions it appears. We also separate it from other large-learning-rate effects. The Maximal Update Parametrization (µP) is widely thought to make the optimal learning rate stay the same as width changes, so it can be tuned on small models and reused on large ones. We show that on some datasets this transfer fails due to the learning rate double descent. As width increases, the optimal learning rate jumps from the low-learning-rate regime to the high-learning-rate regime. As a result, a learning rate tuned on a small proxy model can lead to clearly worse choices at scale.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.