acceptodds
Under review as a conference paper at ICLR 2027

Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining

Abstract

Causal-decoder Transformers form predictions through a hierarchy in which lower layers construct the residual features consumed by upper-layer attention. This paper identifies a GPT pretraining failure mode: upper layers can specialize their Q/K attention patterns too early, before lower-layer representations have stabilized. Temporarily slowing only upper-layer Q/K projections during early training improves final perplexity and downstream accuracy while leaving the rest of the model unchanged. The intervention prevents upper attention from collapsing onto an immature residual basis. In LLaMA-style blocks, the same intervention is largely unnecessary. Through ablations, we trace this difference to multiplicative gated feed-forward networks, which suppress the upstream residual writes that drive the failure. A pathwise analysis connects the two mechanisms: the learning-rate intervention reduces a step-size factor, while gated FFNs reduce a residual-energy factor along the same growth pathway. The results identify upper-layer Q/K timing as a concrete interaction point between decoder architecture and optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.