Spiking, Fast and Slow: The Anatomy of Loss Spikes in Language Model Pretraining
Abstract
Loss spikes in language model pretraining can waste substantial compute and derail otherwise healthy runs. A primary driver of loss spikes is the growth and divergence of output logits in the final language modeling (LM) head. In this work, we divide these training instabilities into two types: fast spikes that abruptly increase loss before recovering, and slow terminal divergences that emerge only after prolonged training, which we call time-bomb spikes. We build on observations that these instabilities stem from fundamental properties of the softmax cross-entropy loss and its interaction with Adam’s preconditioner. We observe that this interaction leads to LM head weight growth that may be harmless under exact arithmetic but can be destabilizing under finite-precision training and propose a simple optimization adjustment that ameliorates the issue. Furthermore, we show that various mitigation techniques from the literature (including our own) largely act on the shared underlying problem of logit growth by reducing quantization error in the LM head. Our work connects fast and time-bomb spikes and an array of prior methodological proposals through a common lens of analysis: controlling last-layer logit growth reduces finite-precision rounding error which is essential for stabilizing long pretraining runs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.