RePair: Mining Multi-Directional Regional Preferences from Imperfect Pairs
Abstract
Post-training plays a crucial role in advancing generative models, yet acquiring reliable reward signals is costly and sparse. Conventional approaches evaluate an entire image with a global reward, implicitly assuming that artifacts are evenly distributed. However, modern text-to-image models usually produce localized failures rather than global degradation. To supervise local errors, recent preference construction heavily rely on manually designed contrasts, which inherently assume that all regional preferences point toward a target image, imposing a restriction we term the unidirectional preference constraint. Instead, we explore a different source of supervision by leveraging the intrinsic stochasticity of text-to-image models, where local successes and failures are often mixed, formalizing this rich signal as multi-directional regional preferences. Building on this insight, we introduce RePair, which transforms cross-generation variance into reliable regional supervision. First, correlated sampling creates candidate images where global structures remain aligned while local differences are preserved. Second, visual grounding directly mines verified regional preferences without requiring artificial editing. Finally, we optimize each regional relation independently using a signed, region-normalized objective, allowing a single image pair to simultaneously provide multiple preference directions. Extensive experiments on FLUX.2 and Ideogram 4 demonstrate that RePair, trained on only 5K automatically mined pairs, achieves substantial improvements across compositional benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.