acceptodds
Under review as a conference paper at ICLR 2027

Binding Crystallization: Diagnosing and Guiding Attribute Assignment in Text-to-Image Diffusion

Abstract

Text-to-image diffusion models can generate high-quality images from natural language descriptions, yet often fail to assign attributes to the intended objects. Existing training-free methods mitigate these errors by guiding attention during sampling, but the dynamics governing when and how attribute assignments can be corrected remain insufficiently understood. Through counterfactual prompt switches and step-local interventions, we identify binding crystallization: across four prompt categories and three diffusion backbones, attribute assignments become resistant to reversal before spatial refinement is complete. Late denoisers still distinguish opposite assignments, but the sampler writes less of that difference into the latent, and comparable latent writes exert weaker downstream effects on ownership. Intermediate-trajectory probes further show that object presence and attribute ownership are distinct control problems. Together, these findings prescribe when, where, and what to guide. CrystalBind is the resulting training-free framework: it sets its intervention window from observed binding transitions, refreshes spatial targets from evolving noun attention maps, and jointly controls presence and ownership through noun excitation, target attraction, and competitor suppression. On T2I-CompBench with SD v1.5, CrystalBind improves the average binding score by 21.3 points over the unguided model and by 3.4 points over CONFORM under a matched budget of 26 guidance updates, while reducing wall-clock time by 28% relative to CONFORM. The same configuration improves SD v2.1 by 14.3 points without retuning. These results show how binding dynamics inform effective and efficient compositional guidance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.