Understanding Learning Dynamics via Symmetry-Aware Spectra
Abstract
Overparameterized neural networks in certain settings have been observed to generalize long after perfect memorization, a phenomenon known as grokking. What drives this transition, and how to detect it during training, remains only partially understood. Common diagnostics are defined on the model weights, but parameter symmetries allow the same function to be represented by weight matrices of different norms and ranks. We study instead the singular-value spectrum of the parameter Jacobian , called the neural spectrum, as a symmetry aware measure of the locally accessible functional directions. We find that across different architectures, grokking is accompanied by spectral polarization: sensitivity becomes concentrated in a small effective subspace. Comparing standard, NTK, and maximal-update (P) parameterizations shows that the timing and form of this restructuring depend strongly on the feature-learning regime. Moreover, regularizing to encourage spectral concentration accelerates grokking, while encouraging a diffuse spectrum deters grokking. Beyond grokking, harder learnable tasks across four supervised-learning benchmarks generally produce broader, higher-rank spectra. These results support the neural spectrum as a common diagnostic of feature-learning dynamics and learned functional complexity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.