When to Trust Inversion: Drift-Aware Preference Optimization for Diffusion Models
Abstract
Direct Preference Optimization (DPO) has emerged as an effective paradigm for aligning text-to-image diffusion models with human preferences. Recent inversion-based methods replace random forward noising with deterministic inversion to provide more precise, image-specific preference supervision. However, as inversion progresses, the resulting latents increasingly drift away from states typically encountered during Gaussian-initialized denoising. This distributional mismatch is particularly pronounced at high-noise timesteps, undermining the effectiveness of inversion-based preference learning. To address this challenge, we propose Inversion-aware Fallback for Preference Optimization (IFPO), a drift-aware preference optimization framework with a dynamic fallback mechanism. IFPO mitigates the impact of inversion drift by adaptively reverting to standard forward noising at high-noise timesteps, while preserving the benefits of inversion-based supervision where it remains useful. Experiments across diffusion and flow-matching backbones demonstrate that IFPO consistently improves text-to-image alignment and visual fidelity over existing preference optimization methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.