acceptodds
Under review as a conference paper at ICLR 2027

ESG: Early Spatial Guidance for Compositional Text-to-Image Generation

Abstract

Text-to-image models often generate the requested objects but fail to arrange them according to the prompt. We present Early Spatial Guidance (ESG), a training-free method for improving spatial composition in rectified flow transformers. ESG reads object locations from native joint attention and updates the latent state to satisfy relative spatial constraints during early sampling. The objective leaves absolute object positions unspecified and handles multiple objects by combining pairwise relations. Normalized updates control intervention strength, while attention-layer calibration adapts the spatial readout to each model. Gradients require only the transformer prefix leading to the selected readout. We analyze the sensitivity of this readout, the interaction of relation gradients, and how interventions propagate through the remaining sampling trajectory. Experiments on SD3 and SD3.5 show improved spatial alignment on public benchmarks and generalization to unseen object pairs. Comparisons with spatial-guidance methods and time-matched native sampling demonstrate a favorable accuracy–cost tradeoff. On controlled multi-object tasks, ESG outperforms the evaluated absolute-centroid objective at matched update budgets, with the clearest gains under repeated guidance. These findings show that early attention-based guidance can coordinate multiple spatial relations at competitive inference cost, without retraining the generator or requiring user-specified layouts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.