RatioDiff: Aligning Diffusion Models via Preference Ratio Matching
Abstract
Diffusion models have achieved remarkable success in text-to-image generation, producing diverse and visually compelling images from natural language prompts. Preference learning further aligns these models with human judgments of image quality and prompt adherence. However, effective preference alignment faces two challenges: weak corrective signals on strongly misranked pairs and difficulty balancing preference improvement with denoising preservation. We introduce RatioDiff, an offline framework centered on ratio diffusion preference learning. Our core contribution is a preference ratio formulation that explicitly accounts for distinct noisy diffusion states through a reference correction. This formulation leads to a tractable Bregman objective that sustains corrective pressure on misranked pairs, providing a principled foundation for diffusion preference optimization. To support this objective, we introduce learned pair weights guided by reference consistency, separating the strength of preference correction from each pair's contribution to training. We further develop adaptive regularization that controls reference proximity and preferred-image denoising relative to the preference update, limiting model drift while preserving denoising behavior. Together, these contributions combine preference ratio matching with explicit control over pair influence and regularization, without additional inference overhead. Experiments on multiple text-to-image backbones and preference benchmarks show that RatioDiff consistently improves human-preference alignment while preserving and enhancing visual quality and text–image consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.