Ternary Preference Optimization for Text-to-Image Generation with Noisy Human Feedback
Abstract
Recent text-to-image alignment methods typically rely on two assumptions about human preference data: (1) annotators always strictly prefer one image over another, and (2) preference labels are noise-free and faithfully reflect underlying human preferences. In practice, both assumptions can be violated. For example, approximately 11% of annotations in Pick-a-Pic are ties, indicating that annotators may consider two images equally preferable. Such ties provide meaningful preference information but are typically discarded by binary preference modeling. Meanwhile, human preference labels are inherently noisy due to individual subjectivity and annotator disagreement. To address these limitations, we propose *Noisy Ternary Preference Optimization* (NTPO), a unified framework that explicitly models three preference outcomes—preferred, rejected, and tied—while accounting for noisy and inconsistent annotations. By preserving tie information and explicitly modeling annotation noise, NTPO provides a more faithful formulation of human preference alignment for text-to-image generation. Experiments across multiple benchmarks demonstrate that NTPO improves both alignment quality and robustness compared with existing binary preference optimization methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.