acceptodds
Under review as a conference paper at ICLR 2027

Understanding Optimization from Activation and Gradient Spectra: A Modded-NanoGPT Study

Abstract

We study the activation covariance and per-sample gradient SVD spectra throughout training in a modded-NanoGPT study. We show that the tail exponents of these spectra reveal differences in internal geometry that the training loss does not. We propose using these spectra as practical predictors of training optimization in two ways. First, we show that the tail exponent of the activation covariance spectrum at an early checkpoint predicts the terminal token efficiency. In controlled experiments in which the runs reach the same terminal loss, it achieves a Spearman correlation of , compared with for the early test loss. Second, we show that along a cumulative chain of modded-NanoGPT interventions, the activation and gradient spectra together distinguish the changes in internal geometry that accompany gains in token efficiency, in throughput, or in both. For example, the MuonUntied transition improves token efficiency by while throughput rises by only , a shift marked primarily by a change in the activation covariance tail. Both signals hold across architectures and for model sizes up to 3.57B parameters. We explain this behavior with an exactly solvable kernel model of next-token prediction on modular arithmetic.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.