acceptodds
Under review as a conference paper at ICLR 2027

LoopLA: Causal State Refinement for Looped Linear Attention

Abstract

Looped Transformers iteratively refine hidden representations at the cost of additional computation. Linear attention can reduce this cost through subquadratic scaling with sequence length. However, the role of its recurrent memory state in looped algorithms has not been systematically studied. We introduce LoopLA, which refines this memory along both token and loop dimensions. To preserve causality without serial token-by-token computation, LoopLA uses a Jacobi-style draft-and-consolidate procedure that approximates state-dependent coefficients for parallel scans. Chunk-wise state commits refresh the memory used by subsequent drafts. We train approximately 1B-parameter Gated DeltaNet models from scratch for 100B tokens. At matched parameter count, LoopLA outperforms baselines without state carry overall and improves the average RULER score by 2.89 points over the no-loop baseline. Under matched FLOPs, it remains comparable, though slightly weaker overall, to the no-loop baseline, with better long-context performance on RULER 2k/4k. This work demonstrates the importance of state refinement along both token and loop dimensions for looped linear attention.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.