acceptodds
Under review as a conference paper at ICLR 2027

Residual Connections with Explicit Embedding Reweighting

Abstract

In a pre-norm Transformer, each token's input embedding enters the residual stream with a fixed coefficient of one. We find that pretrained language models nevertheless reweight the embedding implicitly: across six models from four families, the residual state's projection onto each token's own embedding grows substantially with depth and far exceeds its projection onto other tokens' embeddings. This reweighting is produced indirectly, by attention and MLP layers that, after the first, see the embedding only mixed with earlier updates. To provide a direct path for it, we introduce Explicit Embedding Reweighting (EER), a lightweight module that adds the original embedding with a positive coefficient predicted from the current state. EER keeps the identity shortcut and stores no layer outputs or extra embedding tables. At three scales (356M–900M parameters), EER lowers validation loss by 0.022–0.025 nats with about 0.007% additional parameters. At 900M parameters, it raises average zero-shot accuracy on seven tasks by 1.39 points, is competitive with Attention Residuals, and further improves SATFormer and identity Hyper-Connections. Ablations show that its state-dependent coefficient outperforms fixed, per-layer, and token-only coefficients.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.