Rethinking at Fixed Points in Looped Language Models
Abstract
A looped language model whose recurrent states converge to fixed points can replace its recurrent trajectory with the terminal state, in both training and inference. Near a fixed point, the endpoint can stand in for the whole trajectory and unlock a series of abilities: truncated backpropagation in training, terminal key-value cache reuse that shrinks the cache for a loop- recurrent core at almost no performance loss, distillation for up to faster prefill, and faster post-training updates that compute gradients from saved rollout states instead of replaying the trajectory. To better shape fixed points, we improve two existing techniques: depth sampling, whose depth prior we propose to learn from training signals with constraints; and input injection, which we force orthogonal update so that every recurrence receives the input intact. At M, M, and B parameters, the learned depth prior lowers perplexity under terminal cache reuse by – relative to Huginn's fixed prior; at B, it matches the full-cache downstream accuracy of fixed-depth training. Orthogonal injection lowers perplexity by up to relative to leading injection schemes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.