Measuring the local shape of learning-rate transfer curves
Abstract
Understanding the learning-rate (LR) dependence of trained model performance is essential for better optimization of neural networks, but the appropriate size of the LR depends on learning dynamics, architecture, and their hidden biases, and thus our understanding of this complex dependence remains limited. The behavior of the loss curve as a function of the LR has recently been observed in extensive experiments on LR transfer across different settings, such as model sizes, and it is becoming more important to explore methodologies for quantitatively characterizing the shape of this LR curve. In this work, we provide two relational equations that give a good empirical explanation of the gradient of the LR transfer curve, i.e., the LR gradient. One is a theoretical formula that holds at the edge of stability, and the other is an approximate expression obtained by naively treating this gradient as a hyper-parameter gradient. Using these two relational equations in a complementary manner, we show a lower bound on the gradient of the curve and a zero gradient under specific conditions for normalization layers in theory. The approximate expression also applies to a wide range of experimental settings and allows us to decompose the LR gradient into layer-wise components and quantify the strong contribution of specific layers to the curve shape. Thus, this serves as a foundation for better understanding the internal mechanisms of LR dependence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.