acceptodds
Under review as a conference paper at ICLR 2027

One Model, Many Depths: A Recipe for Elastic-Depth Latent Recurrent Language Models

Abstract

A latent recurrent, or looped, language model applies a shared block of layers repeatedly, so one checkpoint can trade inference compute for accuracy by changing the loop count. At one loop, however, such models typically fall well short of a non-recurrent model of equal inference depth. This one-loop tax makes a separate non-recurrent model the better choice for latency-sensitive serving. We present Elastic Latent Recurrence (ELR), a recipe for training looped language models without this tax. Run at one loop, an ELR model costs the same to serve as a non-recurrent model of the same size and, when the two are trained on the same data, reaches similar downstream accuracy. Each additional loop of the same checkpoint then buys more accuracy for more inference compute. We show that two choices determine whether the one-loop accuracy survives mixed-depth training. The first is a gated update that sets a learned step size along each proposed state revision. The second is mixing loop counts within each optimizer update rather than across updates, which we implement by forcing a rotating subset of data-parallel ranks to run one loop. On a B dense hybrid model trained for 1T tokens and a hybrid mixture-of-experts model with 30B total and 3B active parameters trained for 300B tokens, each alongside a matched baseline, increasing inference depth from one to four loops improves the ten-task average by 7.25 and 4.75 points, respectively. Both models acquire elastic depth late in pretraining. Training at one loop for the first three quarters of pretraining and at mixed depth for the last quarter costs about 1.16 the baseline training FLOPs, and the resulting checkpoint improves at every added loop while giving the best one-loop quality of the schedules we tested. In our hybrid MoE experiments, reusing first-loop expert assignments on later loops removes repeated router evaluation and cuts mixed-depth training update time by 11.7% with no measurable change in downstream accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.