Grow, Don't Prune: Compute-Efficient Training of Diffusion Transformers
Abstract
Diffusion Transformers (DiTs) achieve strong generation quality but are expensive to train and evaluate. Existing dynamic-width methods reduce inference cost by pruning a pretrained dense model, leaving its pretraining cost unchanged. We propose GrowDiT, a dynamic-growth training strategy that starts from a narrow network and progressively activates attention heads and MLP channel groups across diffusion noise levels. A gradient-based need score selects which units to activate under an increasing compute budget. By reducing the FLOPs of early training updates, GrowDiT reallocates the saved computation to additional optimization steps. On ImageNet-256 and FFHQ-256, under a fixed training-FLOP budget and approximately matched inference FLOPs, GrowDiT consistently outperforms fixed-width, random-growth, and dense-training baselines across DiT-S/2, DiT-B/2, and DiT-L/2. On the 130M-parameter DiT-B/2, it improves FID by 7% over fixed-width training while using the same nominal training and inference FLOPs. Matching budgets by wall-clock rather than FLOPs gives a smaller 5.3% on DiT-B/2 and 1.1% on DiT-S/2. GrowDiT also outperforms depth-growth configurations and a DyDiT-inspired train-then-prune baseline. These results suggest that allocating width during training is a promising approach to improving the training efficiency of diffusion transformers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.