acceptodds
Under review as a conference paper at ICLR 2027

ToPO: Token-Oriented Preference Optimization for Text-to-Image Diffusion Models

Abstract

Diffusion models have achieved remarkable success in text-to-image generation, yet aligning their outputs with human preferences remains challenging. Diffusion-DPO learns from image preference pairs, but their global labels do not specify how supervision should be allocated across spatial locations and denoising timesteps. The semantic roles of prompt tokens are also not explicitly considered when assigning these local weights. To address these limitations, we propose Token-Oriented Preference Optimization (ToPO), which builds a token-conditioned spatial–temporal weighting map for each pair. A frozen reference denoiser obtains spatial and temporal weights from local noise-prediction error differences under shared noise, while preferred-branch cross-attention links spatial evidence to content tokens. The map reweights Diffusion-DPO and supports an auxiliary pixel-midpoint ordering loss. It requires neither local preference annotations nor a separately trained reward model. Extensive experiments across multiple preference and compositional evaluations show that ToPO achieves state-of-the-art performance among the compared methods with only 500 training updates, demonstrating the potential of token-conditioned preference allocation for high-quality diffusion alignment within a limited update budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.