acceptodds
Under review as a conference paper at ICLR 2027

Spectral Alignment for Early Detection of Loss Explosion in Language Models

Abstract

Loss explosions can invalidate expensive language-model training runs, yet common scalar monitors such as weight and gradient norms are often ambiguous or late because their scales vary across models and layers. We introduce Spectral Alignment (SA), a data-conditioned and scale-robust monitor of the signed cosine alignment between layer inputs and the principal left singular direction of a weight matrix. Across a mini-batch, mixed-sign SA permits cancellation in the first-order spectral update. We show that a sufficiently one-sided SA distribution, together with compatible downstream gradients, produces positive spectral drift, and derive conditions under which a persistent SA warning precedes a spectral-norm threshold. These results provide a conditional mechanism rather than a universal failure guarantee. In two language-model failure settings, directional alignment collapse emerges substantially before loss explosion and provides a clearer warning than the monitored scalar baselines. SA also selects healthier checkpoints for retraining, leading to lower loss than later pre-failure checkpoints. With configurable monitoring cost through sampled layers, power iteration, and offline computation, SA provides an actionable signal for early intervention and recovery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.