acceptodds
Under review as a conference paper at ICLR 2027

From Front-Loading to Gradual Refinement: Simplex Gating in Transformer Decoders

Abstract

We observe pronounced _front-loading_ of residual-norm growth in transformer decoders: the first block sharply amplifies the residual stream, while subsequent blocks produce much smaller relative increases. Our analysis identifies a structural asymmetry underlying this pattern: pre-normalization standardizes sublayer inputs but leaves residual-update magnitudes unconstrained, giving a write of a given size greater relative influence near the input, where the stream is smallest. Motivated by this asymmetry, we introduce _simplex gating_ to regulate residual writes in both attention and MLP sublayers. An input-dependent scalar gate scales the attention output, while normalized group gates reweight MLP activations before the down-projection. The MLP gates are nonnegative and sum to one for each token, enabling adaptive allocation of a fixed total gate weight across activation groups. Under matched pre-training conditions, simplex gating substantially alleviates front-loading, reducing the first-block residual-norm growth ratio from 9.10× to 1.84×. It also improves pre-training convergence at two model scales, achieving consistently lower training loss. Ablations show that the simplex constraint mitigates compensatory activation growth that can weaken independent sigmoid gates. These findings demonstrate the practical benefits of controlling residual writes and suggest a path from front-loaded residual growth toward more gradual refinement across depth.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.