acceptodds
Under review as a conference paper at ICLR 2027

TADR: TIMESTEP-AWARE DEPTH ROUTING FOR COMPOSITIONAL TEXT-TO-IMAGE GENERATION

Abstract

Diffusion Transformers (DiTs) power today’s strongest text-to-image (T2I) generators, yet inherit one Transformer component unexamined: the residual stream, where every block output accumulates with a fixed unit weight, identically at all denoising timesteps. This rigidity costs diffusion twice. Uniform accumulation dilutes each layer’s contribution and lets prompt semantics fade with depth, degrading precisely the compositional bindings that alignment benchmarks probe; and the aggregation is timestep-agnostic, though denoising demands prompt-driven layout at high noise and texture refinement at low. Learned depth-wise aggregation has recently improved large-scale language-model pre-training, but whether it transfers to pre-trained T2I backbones—the regime practitioners operate in—has remained open. We answer affirmatively, and show that timestep-awareness is what unlocks its potential. We present TADR (Timestep-Aware Depth Routing), which supersedes fixed accumulation in a pre-trained MM-DiT via three components: Block-wise Depth Attention lets each fusion point selectively re-read the block outputs beneath it through softmax attention along depth; a Timestep-Modulated Routing Query injects the denoising timestep into the routing query through a low-rank path; and Zero-Init Identity Gating renders the module an exact identity at initialization, preserving the backbone’s pre-trained competence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.