acceptodds
Under review as a conference paper at ICLR 2027

Spectral Optimization in Residual Networks: Learning Anisotropic Corrections

Abstract

Spectral optimizers such as Muon are often studied through deep linear models, where repeated matrix products create mode-dependent learning rates that become more uneven with depth. Modern pre-norm residual networks operate in a different regime. Their blocks can remain close to the identity map, which makes the classical product-induced factor nearly constant across modes. Meanwhile, the identity path can hide substantial anisotropy in the trainable residual correction. We argue that residual connections therefore shift, rather than remove, the relevant conditioning problem: the identity path stabilizes the full map, while the optimizer must learn a residual correction that has anisotropic structure. Under anisotropic residual-branch geometry in deep linear networks, gradient descent learns correction modes unevenly, while spectral updates balance progress across unlearned directions. Our results suggest that Muon's benefit in deep residual architectures arises because residual parameterization exposes a local spectral correction problem well matched to spectral optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.