acceptodds
Under review as a conference paper at ICLR 2027

Jacobi Recirculation: Parallelizing Recirculation with Fixed-Point Iteration

Abstract

A growing line of work lets a transformer reuse its own deeper representations in later computation, increasing its effective depth without adding parameters. One recent approach is recirculation, which feeds the state of each token at a deeper layer back into a shallower layer, thereby rewriting the key–value cache for later tokens to attend to. However, this token-wise recursion makes processing a known sequence expensive in prompt prefill, evaluation, hyperparameter selection, and the training of its auxiliary mixer. In this work, we propose Jacobi recirculation, which replaces the recursion with a few rounds of parallel updates over all token positions. We prove that the iteration is exact after finitely many rounds, converges geometrically under a local contraction condition, and yields convergent training gradients under additional regularity. Experiments across model families and sizes show that two to three rounds suffice to reproduce sequential recirculation, even though the pretrained models were never trained for this iteration. When used for prompt prefill before generation, it also preserves downstream accuracy. In our experiments, it speeds up single-sequence prefill by more than an order of magnitude and reduces mixer training from hours to minutes. Together, these gains make recirculation substantially cheaper to use and to study. [Code](https://anonymous.4open.science/r/jacobi-recirculation-review-5846/README.md) is available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.