acceptodds
Under review as a conference paper at ICLR 2027

Simulating Training Spikes of Deep Neural Networks by Layer-wise Random Feature Models

Abstract

Intermittent “spikes” in loss, during the pre-training of Large Language Models (LLMs), remains a mysterious phenomenon both theoretically and empirically. In this paper, we show that simple quadratic models with heavy-tailed, power-law distributed inputs can replicate these spikes. We introduce a theoretical framework and simulation algorithm to model per-layer training dynamics of Deep Neural Networks (DNNs), using such quadratic models. In a case study on an Adam hyperparameter configuration (where is set close to ) known to induce spikes in Transformer LMs, our simulation successfully predicts bumps indicating spiky behavior from a given checkpoint. Furthermore, we derive an appropriate schedule from observed metrics, which empirically improves performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.