acceptodds
Under review as a conference paper at ICLR 2027

Lost in Attention: Understanding and Correcting Compositional Failures in Diffusion Models

Abstract

Text-to-image diffusion models remain unreliable at simple compositional requests: objects go missing, counts are wrong, attributes attach to the wrong object and spatial relations are ignored. We trace these failures to how cross-attention allocates its budget: the start token absorbs most of it, the concept named first is favoured, and a concept whose share falls early never appears. The information the text transmits to the image factorises exactly into an allocation and a localization factor, so stating where an object should be cannot make it appear. We propose Mordant, a training-free guidance method that states compositional requirements as constraints on the frozen model's attention, with multipliers found by dual ascent along the sampling trajectory and a relation margin derived from the frame. On Stable Diffusion 1.4, Mordant raises GenEval from 0.45 to 0.70, T2I-CompBench colour binding from 0.37 to 0.69 and 2D spatial relations from 0.13 to 0.39, and VISOR object accuracy from 33.4% to 79.9%; with its constants unchanged it also raises GenEval on SD 2.1 (0.50 to 0.72), SDXL (0.58 to 0.71), SD3-medium (0.70 to 0.83) and SD3.5-medium (0.70 to 0.83).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.