The Mechanism of Training Collapse in Ultra-Deep Transformers
Abstract
Ultra-deep transformers can collapse in training: after a period of successful loss descent a run jumps, often within ten optimizer steps, onto a persistent degenerate plateau, a near-zero predictor for a flow-matching DiT and a near-unigram predictor for a language model. We study this collapse on diffusion transformers of 256–1024 residual sublayers and language models of 256–2048, pushed past their stability boundary by the learning rate, with kernel probes of the residual-stream write matrices, loss profiles along committed AdamW updates, and interventions. Our hypothesis is a condition on one optimizer step, written in output space: the step is harmful when the tangent kernel's gain along the loss gradient, times the learning rate, carries the output past the radius within which the loss still decreases. The tolerable learning rate falls with depth, as a power law over the depths tested; the gain along the gradient rises before any loss symptom and has fallen from its peak on every terminal plateau; and neither the amplified direction nor continued growth of the kernel is required for the instability. GROW, a retraction of each write matrix onto a fixed-Gram set after every AdamW step, is an intervention on this geometry, not a tuned optimizer: keeping the Post-LN forward pass and the zero-initialized write matrices, at four depths from 256 to 2048 residual sublayers it trains through a constant dose that collapses most or all AdamW seeds at the first three, and raising the scale of its fixed factor brings its sample quality to within two FID, inside the run-to-run spread, of a healthy low-rate AdamW reference at every depth.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.