Filling the Slots: Repeated Mentions Get Their Attributes Before Their Instances in MMDiT Text-to-Image Models
Abstract
Text-to-image counting benchmarks phrase multiplicity almost exclusively with numerals ("three bicycles”). We show that enumerating the same objects as repeated mentions ("a bicycle, a bicycle and a bicycle”) instead lowers counting accuracy by 44.3 points in SD3.5-medium and 59.2 points in FLUX.1-dev, a failure that numeral-only evaluation cannot detect. The penalty is not explained by same-category numerosity: repeated mentions lose instances, and cues that distinguish them, such as contrastive determiners and distinct colors, recover far more instances in FLUX.1-dev than in SD3.5-medium, which instead renders the requested colors on too few objects. Scoring counting and color binding separately shows that both models bind colors once the count is right, so the asymmetry lies in instances. Distinct colors raise the separability of the mentions' attention maps early in both models, but this predicts success only in FLUX.1-dev. In SD3.5-medium, the two ways a mention loses its instance have distinct early signatures: an omitted mention's image-to-text attention decays on its own, whereas fused mentions keep their attention but are never given separate regions. Intervening on joint attention shows that restoring early attention to a mention causally restores its color, through the early image layout, but has no specific effect on its instance. A rendering correction built on this mechanism, EMAR, restores missing color bindings and, across six prompt structures, raises counting accuracy by 7–15 points while lowering omission and, wherever a noun repeats, fusion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.