DELAYED DEPTH EXPANSION FOR LOOPED TRANSFORMERS
Abstract
Looped Transformers recently have gained attention for increasing computational depth by repeatedly applying weight-shared Transformers blocks. While improving parameter efficiency, it requires heavy recurrent computation for every token during training. We ask whether this recurrent depth must be paid for throughout the entire training pipeline. To this end, we introduce Delayed Recurrent Depth Expansion (DDE), which first trains a shallower dense Transformer and introduces parameter-shared recurrent depth only later in training. Directly converting a trained dense model to repeated computation causes substantial functional disruption, so we use step-relaxed recurrence to provide a controlled dense-to-recurrent transition. Across 155M and 314M parameter models, DDE substantially reduces cumulative Transformer-block FLOPs at matched validation loss compared to training the corresponding recurrent architecture from scratch. Step relaxation notably reduces the instantaneous loss increase at conversion, with the advantage growing at larger recurrence depths, while remaining competitive with conventional looping when trained from scratch. We further show that the same conversion mechanism can be applied to pretrained language models across multiple model families and billion-parameter scales. These results demonstrate that the recurrent computation used at convergence need not be paid for throughout training, providing a simple route to more compute-efficient training of looped Transformers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.