acceptodds
Under review as a conference paper at ICLR 2027

On the Depth Scaling of Neural Optimizers via L-Smoothness

Abstract

Scaling deep learning models efficiently requires predictable performance transfer across compute regimes. Establishing reliable hyperparameter transfer allows optimal training dynamics and structural choices to be discovered on low-cost proxy models, avoiding catastrophic failures during massive-scale training. Existing depth-transfer analyses are architecture-specific and heuristic: hyperparameter transfer is only justified through toy blocks, and both rely on infinite-depth dynamics, a specific initialization, and feature-learning criteria. Building on -smoothness framework of Xu et al. (2026), we understand hyperparameter transfer as how do the Lipschitz and -smoothness constants of the loss scale with the depth ? We show that inverse-depth residual branch scaling is sufficient to make both constants depth-independent at fixed width, uniformly over the bounded parameter domain and for arbitrary geometries. Because the bounds depend on each residual branch only through its Lipschitz and smoothness constants, the argument applies to general residual networks—including branches of block depth such as transformer and ResNet blocks—without assumptions on initialization or correlations between layer weights. Since the smoothness constant controls the curvature penalty in the descent inequality, our bounds motivate a learning-rate scale without an additional depth-dependent curvature correction for updates bounded in the prescribed norm. This provides a geometric explanation of inverse-depth residual scaling and its application to the operator-norm geometries associated with AdamW, Muon, and MOGA. Learning-rate sweeps on GPT-2 (depths 2–128) and CIFAR-100 ResNets (8–32 residual blocks) confirm that Muon and MOGA exhibit stable learning-rate transfer under these prescriptions, with a common optimal rate across a 64-fold increase in depth.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.