Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
Abstract
Causal-decoder Transformers form predictions through a hierarchy in which lower layers construct the residual features consumed by upper-layer attention. This paper identifies a GPT pretraining failure mode: upper layers can specialize their Q/K attention patterns too early, before lower-layer representations have stabilized. Temporarily slowing only upper-layer Q/K projections during early training improves final perplexity and downstream accuracy while leaving the rest of the model unchanged. The intervention prevents upper attention from collapsing onto an immature residual basis. In LLaMA-style blocks, the same intervention is largely unnecessary. Through ablations, we trace this difference to multiplicative gated feed-forward networks, which suppress the upstream residual writes that drive the failure. A pathwise analysis connects the two mechanisms: the learning-rate intervention reduces a step-size factor, while gated FFNs reduce a residual-energy factor along the same growth pathway. The results identify upper-layer Q/K timing as a concrete interaction point between decoder architecture and optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.