acceptodds
Under review as a conference paper at ICLR 2027

Seeds Teach Each Other Through Prompt Tokens: Reward-Free Compositional Text-to-Image Self-Improvement

Abstract

When a text-to-image model samples one compositional prompt with several seeds, different samples get different requirements right: one has the correct count, another the correct spatial relation. Reward and preference fine-tuning score each image as a whole, so they raise the probability of the best complete image and discard the correct parts of the others, and the parts cannot be mixed in pixel space because each seed lays out its objects differently. We observe that joint-attention transformers already provide positions that correspond across seeds: the prompt tokens. After one joint-attention block, the state of each prompt token has absorbed what its own sample rendered for that word, while its position denotes the same word in every sample. VoTEx (Visual-to-Text Exchange) uses this correspondence during training. For samples of a prompt, the next block lets every image query of a target sample attend to the prompt-token states of all samples, with a logit shift that keeps attention unchanged when the samples coincide. The resulting prediction is a dense target that we distill into the ordinary single-sample model, so architecture and inference cost stay the same. We prove that query-wise reading covers requirements with samples under a fixed evidence margin, whereas whole-image selection needs exponentially many, and that the shift is the unique scale-preserving correction. On SD3.5 Medium, FLUX.1-dev, and Qwen-Image-2512, VoTEx improves six compositional benchmarks by 2.2 to 8.9 points without any reward model, with all 18 run-level 95% intervals above zero. Filling the group with copies or other prompts removes at least 90% of the gain, and masking the sample that rendered a missed requirement correctly lowers accuracy on it four times more than masking an unrelated sample.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.