TADR: TIMESTEP-AWARE DEPTH ROUTING FOR COMPOSITIONAL TEXT-TO-IMAGE GENERATION
Abstract
Diffusion Transformers (DiTs) power today’s strongest text-to-image (T2I) generators, yet inherit one Transformer component unexamined: the residual stream, where every block output accumulates with a fixed unit weight, identically at all denoising timesteps. This rigidity costs diffusion twice. Uniform accumulation dilutes each layer’s contribution and lets prompt semantics fade with depth, degrading precisely the compositional bindings that alignment benchmarks probe; and the aggregation is timestep-agnostic, though denoising demands prompt-driven layout at high noise and texture refinement at low. Learned depth-wise aggregation has recently improved large-scale language-model pre-training, but whether it transfers to pre-trained T2I backbones—the regime practitioners operate in—has remained open. We answer affirmatively, and show that timestep-awareness is what unlocks its potential. We present TADR (Timestep-Aware Depth Routing), which supersedes fixed accumulation in a pre-trained MM-DiT via three components: Block-wise Depth Attention lets each fusion point selectively re-read the block outputs beneath it through softmax attention along depth; a Timestep-Modulated Routing Query injects the denoising timestep into the routing query through a low-rank path; and Zero-Init Identity Gating renders the module an exact identity at initialization, preserving the backbone’s pre-trained competence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.