acceptodds
Under review as a conference paper at ICLR 2027

How to Loop Your DiT: Retrofitting Loops on Image Generation Models

Abstract

Scaling inference-time compute has emerged as a prominent avenue for improving neural network performance. Looping models, where transformers are recursively applied to refine predictions, offer a particularly attractive route by allowing a user to control the effective depth of the model at inference time. In the context of image generation, looped transformers trained from scratch offer not only inference-time scaling, but can often match the performance of larger models using significantly fewer parameters. However, training looped models is expensive: an -loop model requires roughly times the FLOPs in both training and inference, negating much of the efficiency advantage. We tackle these limitations by retrofitting loops on top of pretrained models. We focus on the family of ImageNet-scale diffusion models and show that naively looping them often leads to limited improvement in performance. Instead, we draw on recent recursive reasoning literature and present a recipe that uses the looped model to refine a set of context tokens, rather than the main answer stream. We show that this approach can be applied on top of pre-trained models to recover performance competitive with from-scratch looping at a fraction of the training cost. Moreover, models trained in this manner exhibit favorable test-time scaling, and can be easily distilled to lower loop counts with negligible quality loss. As a concrete example, with only a third of the training budget, we retrofit a FLUX.2-style XL/2 model to a -loop variant that matches a larger feed-forward model, and distill it down to loops with little degradation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.