Tracing Compositional Instability: Temporal Dynamics of Attribute-Object Correspondence in Diffusion Transformers
Abstract
Modern text-to-image diffusion models achieve remarkable visual fidelity, yet can fail even on simple compositions in which two objects must each realize their respective attribute. We study how such failures emerge in multimodal diffu- sion transformers (MMDiTs), where text and image representations are jointly updated throughout generation. Evaluation based only on the final generated im- age cannot reveal how these failures develop internally: is the requested com- position already incorrect from the beginning, or does it change as denoising proceeds? This requires controlled generations in which the same semantic con- stituents lead to different compositional outcomes. We therefore introduce a con- trolled compositional benchmark that systematically varies attribute-object com- positions and the conditions under which different bindings compete. Analysing token-representation geometry across blocks and denoising time, we trace how attribute-object bindings evolve internally. For the state-of-the-art MMDiT mod- els SD3.5 and FLUX.1-dev, we find that the requested binding can be supported at earlier stages of the generative computation but become unstable as information is transformed across model depth and denoising time. Finally, we show that these internal signals are functionally actionable and introduce a test-time steering ap- proach toward the requested binding that reduces generation failure. Together, our results characterize compositional binding not as a fixed property of the prompt representation, but as dynamically evolving throughout the generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.