acceptodds
Under review as a conference paper at ICLR 2027

Dual-Mod Attention: Token Recurrence Beats Looping

Abstract

Transformers have shown incredible gains in reasoning performance by scaling parameter count and test time compute, but together they create an immense increase in computation budget and context lengths. Latent-thinking approaches grow computation without emitting tokens, along two paths: looped transformers, which iterate a weight-tied block, and recurrent hybrids, which carry a compact evolving state. Both of these methods have their own failure modes. Looped transformers can solve some problems outside of , but their per-iteration update shrinks and accuracy saturates, after which further iterations add nothing. Their reach is capped by the training-time iteration count and length, and we find that extra test-time iterations do not extend it. Past the training length, looping rapidly collapses to chance and stays within a transformer's length-generalization boundary. Recurrent methods are better at length generalization and also solve tasks outside of , but carry a fixed state that provably bounds them as the number of tracked states grows. Linear recurrences commonly used in hybrids cannot perform -complete state tracking: diagonal ones remain in outright (Illusion of State), and single-step delta-rule ones confine their transition spectra to [0,1]. We present a new path to scaling computation using a non-linear recurrence within softmax attention. Our method uses an append-only -tape where the entries are computed recurrently from the tape instead of raw tokens. We call this class of architectures token recurrences. Our arch is an exact superset of transformers whose chunk sweeps have a unique fixed point, the sequential recurrence itself, reached in finitely many steps, and whose state moves with every token. Since AR decoding is already sequential, the recurrence adds no decoding cost. Empirically, a single layer of our arch solves instances of the P-complete circuit value problem up to depth 1024. It also solves the -complete word problem, and on a variant with interleaved registers tracks up to 1024 registers at full depth, where single-pass attention and fixed-state baselines stop by 32 and 128 registers. Token recurrences cross the boundary that loops cannot while keeping a growing cache, and ours allows chunk-parallel training at a fraction of a layer per sweep. On language modeling tasks, it beats parameter matched baselines on loss and is competitive with SOTA arches in each eval suite.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.