How Nonlinear Dynamics Shape Neural Scaling Laws and the Hyperparameter Landscape
Abstract
How does evolving curvature constrain the scaling of neural network training? Solvable models of neural scaling laws are mostly linear, with a fixed curvature, so they cannot capture how the growing curvature of real networks interacts with minibatch noise or the hyperparameter phenomena that depend on it. We study a minimal nonlinear extension, a two-layer linear network trained with SGD on a power-law task, and derive why and how fast its curvature grows, as feature learning ties the scale of the kernel to the norm of the function learned so far. The growing kernel amplifies minibatch noise until the noise level self-stabilizes at a fraction of the stochastic edge of stability set by the task, at any learning rate and even without finite-step-size effects. This feedback sets a speed limit on training, so the accelerated exponent that prior infinite-batch analyses found for feature learning is only a transient, after which the loss respects an information-theoretic bound that infinite-batch theory violates and the compute-optimal exponent returns to its lazy value. The same feedback explains why the loss is insensitive to the learning rate over a range that widens during training, why the effective learning rate follows the linear scaling rule up to a critical batch size without tuning, and why annealing yields large gains that tolerate long decays, effects that in linear models hinge on the learning rate chosen by the practitioner. SGD experiments across tasks, model sizes, batch sizes and learning rates confirm our quantitative predictions, which connect scaling laws, curvature dynamics and the hyperparameter landscape within a single tractable model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.