LayerLoop: Stage Granularity in Recurrent-Depth Language Models
Abstract
Recurrent-depth Transformers reuse stored blocks to increase executed depth, but the number of blocks repeated together is another architectural choice. We introduce LayerLoop, which divides the recurrent middle into consecutive multi-block stages and repeats each stage before advancing.This stage-partitioned design spans the space between repeating each block before advancing (blockwise recurrence) and repeating the entire middle region as one unit (whole-middle recurrence).We train and compare stage groupings at matched stored and executed depths, repeat count, and processed tokens. In the four-repeat comparison, four-block stages score 71.2 on a 24-component diagnostic suite of reference-relative percentile ranks, versus 55.4 for blockwise and 57.5 for whole-middle recurrence. Blockwise recurrence instead has the numerically lowest measured next-token prediction loss, showing that stage granularity changes the trained model's quality profile even at fixed depths. At similar parameter storage and processed tokens, the four-block-stage model also leads the tested dense and recurrent baselines on the suite. Across two tested widths, its stored parameters participate more broadly than those of matched dense models. Without predefined roles, its stages develop distinct, repeatable responses when their repeated blocks are bypassed. LayerLoop thus turns stage granularity into an explicit design choice, with trained-model behavior that stored and executed depth counts alone do not describe.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.