acceptodds
Under review as a conference paper at ICLR 2027

Delta Attention Residuals

Abstract

Attention Residuals (AttnRes) replace the standard residual connection with learned softmax attention over previous sublayer outputs, and the routed mixture replaces the residual stream as the next sublayer's input. We find that this replacement limits AttnRes at scale: it improves on standard residuals at 220M but is 6.9% worse in perplexity at 1044M and 6.6% worse at 7.6B, and the variant without any residual stream degrades most. We propose Delta Attention Residuals, whose core change is the update rule: they keep the residual stream and add a softmax-routed mixture of deltas (v_i = h_i+1 - h_i) to it. Under an additive update, the natural alternative source, cumulative hidden states, is highly redundant (adjacent cosine 0.93-0.97, R^2 up to 0.99 in deep layers) and gives higher perplexity than deltas on two architecture families. Across 220M-7.6B, Delta Attention Residuals reduce validation perplexity by 1.7-8.2% relative to standard residuals, and a zero-initialized gate on the added term supports fine-tuning from pretrained checkpoints. Code is in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.