One Model, Many Depths: A Recipe for Elastic-Depth Latent Recurrent Language Models
Abstract
A latent recurrent, or looped, language model applies a shared block of layers repeatedly, so one checkpoint can trade inference compute for accuracy by changing the loop count. At one loop, however, such models typically fall well short of a non-recurrent model of equal inference depth. This one-loop tax makes a separate non-recurrent model the better choice for latency-sensitive serving. We present Elastic Latent Recurrence (ELR), a recipe for training looped language models without this tax. Run at one loop, an ELR model costs the same to serve as a non-recurrent model of the same size and, when the two are trained on the same data, reaches similar downstream accuracy. Each additional loop of the same checkpoint then buys more accuracy for more inference compute. We show that two choices determine whether the one-loop accuracy survives mixed-depth training. The first is a gated update that sets a learned step size along each proposed state revision. The second is mixing loop counts within each optimizer update rather than across updates, which we implement by forcing a rotating subset of data-parallel ranks to run one loop. On a B dense hybrid model trained for 1T tokens and a hybrid mixture-of-experts model with 30B total and 3B active parameters trained for 300B tokens, each alongside a matched baseline, increasing inference depth from one to four loops improves the ten-task average by 7.25 and 4.75 points, respectively. Both models acquire elastic depth late in pretraining. Training at one loop for the first three quarters of pretraining and at mixed depth for the last quarter costs about 1.16 the baseline training FLOPs, and the resulting checkpoint improves at every added loop while giving the best one-loop quality of the schedules we tested. In our hybrid MoE experiments, reusing first-loop expert assignments on later loops removes repeated router evaluation and cuts mixed-depth training update time by 11.7% with no measurable change in downstream accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.