acceptodds
Under review as a conference paper at ICLR 2027

LayerLoop: Stage Granularity in Recurrent-Depth Language Models

Abstract

Recurrent-depth Transformers reuse stored blocks to increase executed depth, but the number of blocks repeated together is another architectural choice. We introduce LayerLoop, which divides the recurrent middle into consecutive multi-block stages and repeats each stage before advancing.This stage-partitioned design spans the space between repeating each block before advancing (blockwise recurrence) and repeating the entire middle region as one unit (whole-middle recurrence).We train and compare stage groupings at matched stored and executed depths, repeat count, and processed tokens. In the four-repeat comparison, four-block stages score 71.2 on a 24-component diagnostic suite of reference-relative percentile ranks, versus 55.4 for blockwise and 57.5 for whole-middle recurrence. Blockwise recurrence instead has the numerically lowest measured next-token prediction loss, showing that stage granularity changes the trained model's quality profile even at fixed depths. At similar parameter storage and processed tokens, the four-block-stage model also leads the tested dense and recurrent baselines on the suite. Across two tested widths, its stored parameters participate more broadly than those of matched dense models. Without predefined roles, its stages develop distinct, repeatable responses when their repeated blocks are bypassed. LayerLoop thus turns stage granularity into an explicit design choice, with trained-model behavior that stored and executed depth counts alone do not describe.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.