Reading and Writing Residual States for Diffusion Transformers
Abstract
Diffusion Transformers have advanced mainly through improved backbones, representations, and objectives, while the residual pathway that carries information across depth remains largely fixed. This work investigates an orthogonal question, i.e., how should a denoiser read from and write to multiple persistent residual states? We present a systematic study of this design in diffusion Transformers, using Gated Residual as a starting point. Controlled experiments show that adaptive access improves generation and that the scale of the mixed feature substantially affects its effectiveness. We introduce a direction-calibrated reader that preserves the direction selected by adaptive gates and matches the mixture's magnitude to an ungated average of normalized states. For writing, we find that modulating the update component orthogonal to each stored state outperforms modulating the entire update. This motivates a noise-conditioned tangent writer that controls this component according to the noise level. Both designs retain a single forward of each attention and MLP sublayer across states. The resulting method improves both pixel- and latent-space diffusion models and remains compatible with representation alignment methods. Code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.