acceptodds
Under review as a conference paper at ICLR 2027

Loop Attention, not the Block

Abstract

Recurrent-depth Transformers increase depth by repeating shared parameters, but most loop an entire block. That means recomputing both self-attention and feed-forward layers at every step, which inflates compute and often hurts language modeling. We move recurrence entirely inside self-attention. We introduce Gated Looped Attention (G-Loop), which keeps the feed-forward network outside the loop and repeatedly remixes the input values with the evolving state. Learned token- and channel-wise gates decide how much of each source to write at each step, trained end-to-end with backpropagation through time. As a result, a standard encoder Transformer handles both algorithmic reasoning and general language tasks without specialized routing and with modest parameter overhead. At 32 recurrent steps, G-Loop cuts inference FLOPs by up to 94.7% and memory by up to 73.9% compared to HRM and TRM-Att. It uses about 19% fewer parameters than TRM-Att to reach 85.9% on Maze-Hard and 72.5% on Sudoku-Extreme, beating its reported Maze score and closely approaching its Sudoku score. Under our cost protocol, the newer recurrent reasoners that score higher spend 57× to 4,900× as many inference FLOPs. Added to DeBERTa-v3-base, G-Loop improves the six-task GLUE development average by 1.25 points and SQuAD 2.0 F1 by 0.9 points over the baseline. Looping attention alone slashes compute, and gating the value pathway preserves the expressivity needed to reason.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.