acceptodds
Under review as a conference paper at ICLR 2027

Understanding and Mitigating AdamW Loss Spikes

Abstract

Loss spikes can disrupt AdamW training despite gradient clipping. Mismatched first- and second-moment time scales can let momentum outpace its normalization, amplifying the adaptive update (). We study how its magnitude and allocation affect immediate loss. Across three GPT-2 355M runs, at the unit threshol , the median fraction of coordinates newly crossing the threshold is 0.571% at 19 captured loss-spike onset states versus 0.0033% at nine ordinary states (over ). Globally shrinking the adaptive update lowers immediate loss at all 19 states, while selectively limiting new tail entrants improves further at equal norm. Tensor-relative history helps allocate attenuation, and replay favors sustained over onset-only control. These findings motivate , a lightweight yet effective combination of coordinate clipping and a tensor-relative growth constraint, with negligible optimizer-state overhead and no extra forward or backward passes. Across three GPT-2 774M runs per configuration, it reduces post-warmup spike episodes from 17 to zero and improves mean validation loss from 2.916 to 2.856. Half-rate and global norm-matched controls retain spikes despite similar validation quality. Dual control also achieves zero post-warmup spikes in the tested larger NanoGPT and Qwen3 models, supporting transfer across architectures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.