The Spectral Lazy Regime: Spectral Norms and Generalization in Continual Learning
Abstract
Loss of plasticity, the progressive degradation of a neural network's ability to learn, has been observed in several continual learning settings. This phenomenon can be split into two related ideas: loss of trainability, when a network loses its ability to reduce the training loss, and loss of generalization, when a network can still be trained but its generalization abilities are impaired. We focus on generalization and propose a mechanism for when this can arise: entering the spectral lazy regime. This learning regime is signaled by large increases in the spectral norm of the parameter matrices. When this occurs, the learning dynamics more closely resemble training in the lazy (kernel) regime—representation learning is inhibited, leading to poor generalization. We support this finding by establishing a novel theoretical connection between large spectral norms and the lazy regime. We empirically validate that training a network with constraints on the spectral norm of its weight matrices avoids loss of generalization. Through careful investigations into other norm constraints, we conclude that other choices can be useful but only the spectral norm leads to the radius sensitivity curve being invariant to the width of the matrix. Furthermore, to enable finer control over the spectral norm, we propose two new parameterization of weight matrices, the singular value parameterization and the spectral norm parameterization. We show that these approaches naturally reduce the spectral norm growth compared to standard parameter matrices and can mitigate loss of generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.