acceptodds
Under review as a conference paper at ICLR 2027

Predicting characteristic learning rate based on second-moment-shaped stochastic mobility

Abstract

Neural networks are trained using local gradients with finite step size, making the learning rate central to balancing exploration in the loss landscape and stability in training dynamics. To effectively navigate a highly heterogeneous loss landscape, adaptive training is widely used in which training steps are further adjusted adaptively using the second moment. However, the selection of learning rate remains empirical and often costly. Here we derive a statistical physics prediction for this parameter. We find that the competition between stochastic exploration and curvature-induced confinement in training can be quantitatively defined, so that there exists an optimal characteristic . We estimate from a few mini-batches at a cost of 31 gradient-equivalent steps, rather than selecting it by comparing final performance across multiple training runs. Across deep neural networks and large-scale language models, the predicted lies near the generalization optimum and closely matches the best learning rate found by grid search. Moreover, we provide a measurable criterion for identifying the predictor's applicability regime and show that it attains the highest development accuracy among seven inexpensive selection baselines. Therefore, learning rate selection can be transformed from a costly empirical tuning procedure into an intrinsic dynamical prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.