acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Loop: From Latent Dynamics to Performant Recurrent Models

Abstract

Recurrent Transformers offer a parameter-efficient way to increase latent computation through repeated application of shared modules. However, it remains unclear how to design better-performing recurrent architectures, train them effectively and efficiently, and assess how broadly their gains transfer. To address these questions, we examine HRM-Text as a representative structured recurrent model and find that repeated -module updates refine the state toward a narrow recurrent region, while the module applies a distinct finite transformation. This motivates , a simpler one-stream recurrent family that preserves this functional separation. Under the original training recipe, matches HRM-Text at matched trainable and executed depth; with our new recipe, its eight-benchmark mean improves by 4.8–5.7 points across both scales and surpasses recipe-matched HRM-Text. We further show that truncated backward windows preserve useful gradient directions, while shorter differentiable paths and parameter sharing reduce 60B-token training time by more than relative to matched-depth standard Transformers in our 8×H100 production setting. We finally examine how these gains extend beyond the training mixture. They persist on source-matched evaluations and controlled arithmetic shifts, while broader transfer remains mixed under the current breadth and scale of pretraining. As a complementary adaptability test, fine-tuning on a safety domain absent from pretraining substantially improves safety performance while preserving most general-task utility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.