DeltaPO: Counterfactually Verified Preference Optimization for Text-to-Image Generation
Abstract
Preference optimization for text-to-image generation typically relies on image-level labels that indicate which output is preferred, but not why. When candidate images differ in several respects, such supervision can mix better prompt satisfaction with incidental visual differences. We introduce DeltaPO, a preference optimization framework that tests the semantic basis of a preference through targeted image interventions. DeltaPO decomposes a prompt into verifiable requirements and identifies a semantic delta between matched candidates, together with the image regions that support it. It then tests this delta by repairing the corresponding violation in the rejected image and introducing the same violation into the preferred image. A delta is retained only when at least one intervention is valid and every valid intervention weakens the original preference while preserving non-target requirements and visual quality. The verified delta is then used to select training pairs and localize preference optimization to the implicated regions, while training remains conditioned on the original prompt and original image pair. Using SDXL as the base generator, DeltaPO improves compositional alignment on GenEval, T2I-CompBench, DPG-Bench, and GenEval 2, including gains from 0.53 to 0.64 on GenEval and from 73.38 to 79.97 on DPG-Bench. Ablations show that counterfactual verification provides gains beyond edit validity alone, while localized optimization further improves performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.