StairDrop: Hierarchical Token Dropping for Diffusion Transformers
Abstract
Recent diffusion transformers have demonstrated that substantial training compute can be saved by allowing only a subset of tokens to traverse intermediate blocks while preserving competitive sample quality. Yet existing token-sparse designs usually treat the resulting sparsity pattern as a fixed computational shortcut, leaving open how information should be organized across depth. Inspired by U-Net's progressive contraction and expansion, we introduce Stairdrop, a hierarchical token-dropping schedule that progressively reduces and restores participation, with skip fusion at each restoration and depth allocated across the resulting stages. The hierarchy is used during training; at inference, all tokens traverse the full backbone with the fusion layers retained. Controlled experiments on SiT-B/2 and SiT-XL/2 show that intermediate participation improves generation and that depth allocation matters. On ImageNet with guided sampling, at half of the training compute of the strongest token-dropping system, Stairdrop matches its classifier-free-guidance FID-50K of 1.96, and at equal compute it reaches 1.47 against its best reported 1.55. These results support hierarchical token participation as an effective design for efficient diffusion transformer training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.