acceptodds
Under review as a conference paper at ICLR 2027

DELAYED DEPTH EXPANSION FOR LOOPED TRANSFORMERS

Abstract

Looped Transformers recently have gained attention for increasing computational depth by repeatedly applying weight-shared Transformers blocks. While improving parameter efficiency, it requires heavy recurrent computation for every token during training. We ask whether this recurrent depth must be paid for throughout the entire training pipeline. To this end, we introduce Delayed Recurrent Depth Expansion (DDE), which first trains a shallower dense Transformer and introduces parameter-shared recurrent depth only later in training. Directly converting a trained dense model to repeated computation causes substantial functional disruption, so we use step-relaxed recurrence to provide a controlled dense-to-recurrent transition. Across 155M and 314M parameter models, DDE substantially reduces cumulative Transformer-block FLOPs at matched validation loss compared to training the corresponding recurrent architecture from scratch. Step relaxation notably reduces the instantaneous loss increase at conversion, with the advantage growing at larger recurrence depths, while remaining competitive with conventional looping when trained from scratch. We further show that the same conversion mechanism can be applied to pretrained language models across multiple model families and billion-parameter scales. These results demonstrate that the recurrent computation used at convergence need not be paid for throughout training, providing a simple route to more compute-efficient training of looped Transformers.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.