How Query–Key Initialization Shapes Gradients and Training in Transformers
Abstract
We study how query and key (QK) initialization shapes gradient propagation and training in Transformers. We derive an explicit gradient-norm recurrence for softmax attention and Pre-Norm residual blocks, revealing an asymmetry between large and small initialization scales. Large initialization can cause gradient explosion: the expected input-gradient gain diverges in the long-sequence limit. For normalized, weakly correlated inputs, this yields as a practical upper reference for , where QK weight entries have variance , is the input dimension, and is the sequence length. Small, nonzero initialization preserves a nonvanishing gradient-to-weight ratio, while gradient normalization in Adam and Muon permits substantial updates despite nearly uniform attention. Experiments with -layer attention stacks and 124M-parameter language models support the analysis, with AdamW and Muon exhibiting a broad low-loss plateau down to . These findings motivate a practical recipe: favor small, nonzero QK initialization; and for broader searches, stay below the upper reference .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.