R-DiT: Recursive Diffusion Transformer
Abstract
Diffusion models have primarily scaled along two axes: increasing the number of model estimations and increasing the capacity of the denoiser. The former exhibits diminishing returns as sampling depth grows, while the latter improves generation at the cost of increasingly large parameter and memory footprints. We introduce Recursive Diffusion Transformer (R-DiT), which introduces recursive depth as an additional compute-scaling axis. R-DiT repeatedly refines its hidden representation through shared Transformer blocks, increasing within-step computational depth without proportionally increasing the number of unique parameters. To make shared computation effective, we introduce recursion modulation, which conditions the denoiser jointly on diffusion time and recursion time. We further propose adaptive recursion, allocating different amounts of recursive computation across the denoising trajectory under a fixed compute budget. Across image, language, and molecular diffusion, we show that R-DiT significantly improves the quality-compute-parameter trade-off over standard diffusion Transformers. Our results establish recursive computation as a complementary scaling axis alongside sampling depth and model size in training diffusion models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.