acceptodds
Under review as a conference paper at ICLR 2027

How Query–Key Initialization Shapes Gradients and Training in Transformers

Abstract

We study how query and key (QK) initialization shapes gradient propagation and training in Transformers. We derive an explicit gradient-norm recurrence for softmax attention and Pre-Norm residual blocks, revealing an asymmetry between large and small initialization scales. Large initialization can cause gradient explosion: the expected input-gradient gain diverges in the long-sequence limit. For normalized, weakly correlated inputs, this yields as a practical upper reference for , where QK weight entries have variance , is the input dimension, and is the sequence length. Small, nonzero initialization preserves a nonvanishing gradient-to-weight ratio, while gradient normalization in Adam and Muon permits substantial updates despite nearly uniform attention. Experiments with -layer attention stacks and 124M-parameter language models support the analysis, with AdamW and Muon exhibiting a broad low-loss plateau down to . These findings motivate a practical recipe: favor small, nonzero QK initialization; and for broader searches, stay below the upper reference .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.