From Front-Loading to Gradual Refinement: Simplex Gating in Transformer Decoders
Abstract
We observe pronounced _front-loading_ of residual-norm growth in transformer decoders: the first block sharply amplifies the residual stream, while subsequent blocks produce much smaller relative increases. Our analysis identifies a structural asymmetry underlying this pattern: pre-normalization standardizes sublayer inputs but leaves residual-update magnitudes unconstrained, giving a write of a given size greater relative influence near the input, where the stream is smallest. Motivated by this asymmetry, we introduce _simplex gating_ to regulate residual writes in both attention and MLP sublayers. An input-dependent scalar gate scales the attention output, while normalized group gates reweight MLP activations before the down-projection. The MLP gates are nonnegative and sum to one for each token, enabling adaptive allocation of a fixed total gate weight across activation groups. Under matched pre-training conditions, simplex gating substantially alleviates front-loading, reducing the first-block residual-norm growth ratio from 9.10× to 1.84×. It also improves pre-training convergence at two model scales, achieving consistently lower training loss. Ablations show that the simplex constraint mitigates compensatory activation growth that can weaken independent sigmoid gates. These findings demonstrate the practical benefits of controlling residual writes and suggest a path from front-loaded residual growth toward more gradual refinement across depth.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.