acceptodds
Under review as a conference paper at ICLR 2027

On the Effectiveness and Inductive Bias of Block Grouping in Attention Residuals

Abstract

Attention Residuals (AttnRes) make residual history addressable, raising a natural follow-up question: what should constitute a unit of retrieval? Full AttnRes exposes every sublayer output independently, whereas Block AttnRes groups outputs and forces them to share a readout coefficient. We find that this restriction can improve both learning and efficiency. On FineWeb-Edu, Block outperforms Full for all three seeds at 124M parameters and both seeds at 350M under our standard recipe, reversing the original AttnRes ordering. To separate candidate count from candidate construction, we compare Block with a local non-aligned control matching its final group count and sizes, with similar per-reader candidate counts. Block remains better across the reported comparisons, showing that candidate-count reduction alone is insufficient to explain the gain. Learning-rate sweeps at both scales and a fourfold training-budget range at 124M show that the Block-Full ordering survives retuning, even when retuning restores Full's advantage over standard residuals. With a fused Triton implementation, Block uses approximately 14% and 26% less GPU time per token than Full at 124M and 350M. These results show that residual-source grouping is not merely an efficiency choice but a meaningful inductive bias: how residual history is partitioned can matter as much as how finely it is accessed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.