Understanding and Mitigating AdamW Loss Spikes
Abstract
Loss spikes can disrupt AdamW training despite gradient clipping. Mismatched first- and second-moment time scales can let momentum outpace its normalization, amplifying the adaptive update (). We study how its magnitude and allocation affect immediate loss. Across three GPT-2 355M runs, at the unit threshol , the median fraction of coordinates newly crossing the threshold is 0.571% at 19 captured loss-spike onset states versus 0.0033% at nine ordinary states (over ). Globally shrinking the adaptive update lowers immediate loss at all 19 states, while selectively limiting new tail entrants improves further at equal norm. Tensor-relative history helps allocate attenuation, and replay favors sustained over onset-only control. These findings motivate , a lightweight yet effective combination of coordinate clipping and a tensor-relative growth constraint, with negligible optimizer-state overhead and no extra forward or backward passes. Across three GPT-2 774M runs per configuration, it reduces post-warmup spike episodes from 17 to zero and improves mean validation loss from 2.916 to 2.856. Half-rate and global norm-matched controls retain spikes despite similar validation quality. Dual control also achieves zero post-warmup spikes in the tested larger NanoGPT and Qwen3 models, supporting transfer across architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.