Same Spectral Bulk, Different Stable Steps: State Memory in Wide Neural Networks
Abstract
A common approach to understanding neural network training is to describe directional sensitivities through eigenvalue distributions and use this description to assess stable learning rates. This seems natural: the distribution records how many directions are relatively flat and how many are sensitive. We show that even this overall description can be insufficient to determine local training stability. We construct wide neural networks with the same limiting eigenvalue distribution and all fixed-order normalized spectral moments, yet the same learning rate is stable for one and unstable for the other. Stability can be constrained by a single exceptional direction whose contribution to the distribution vanishes as width grows. Our analysis shows that matching aggregate statistics does not match the critical directions: the network computation can selectively amplify a few directions that determine stability while leaving the limiting distribution unchanged. We characterize when these directions emerge and when they restrict the learning rate in Gaussian networks with two nonlinear layers, for a single input near zero training error with all weights trained. We further give explicit conditions under which a trainable output layer preserves the difference or averages it away. Experiments confirm that networks with nearly identical spectral bulk can admit different stable learning rates, and that a few forward-pass statistics capturing the exceptional directions improve prediction of this difference. These findings can help guide neural network training by identifying the information needed beyond aggregate spectral descriptions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.