Allocating Recurrent Compute in Looped Language Models
Abstract
Looped language models improve reasoning and knowledge manipulation by applying shared computation repeatedly. Existing systems usually repeat an entire layer stack, although a mixer and a dense feed-forward network (FFN) perform different operations and have different costs. We ask a narrower question: what should loop? We view recurrence as iterative refinement of a hidden representation and separate two effects that full-block recurrence couples: repeated context-dependent state updates in the mixer and repeated position-wise transformation in the FFN. This view motivates MixerLoop, which repeats each Gated DeltaNet mixer while applying its dense FFN once. Across matched pretraining experiments, MixerLoop preserves the performance benefits of full-block recurrence while reducing backbone projection compute by 43%, and consistently outperforms NoLoop. We observe that successive mixer applications continue to reduce next-token cross-entropy. MixerLoop also creates an opportunity during inference because mixer weights are reused across recurrent steps. We implement a custom FPGA accelerator on a KV260 that keeps these weights on chip and overlaps additional mixer computation with the transfer of the larger FFN weights. Using the same hardware and compute datapath, NoLoop and four-step MixerLoop both achieve approximately 152 tokens/s. In memory-bound decoding, MixerLoop therefore turns weight-transfer time into useful recurrent computation with nearly unchanged latency. Together, these results identify the token mixer as an effective recurrence boundary, retaining the benefits of recurrent depth while placing additional computation where it can be efficiently reused by hardware during inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.