acceptodds
Under review as a conference paper at ICLR 2027

Aligned, but More Capable? Evaluating Post-Trained Text-to-Image Models Beyond Preference Metrics

Abstract

Preference alignment is widely used to improve text-to-image generation by making model outputs better match human preferences. Yet it remains unclear whether such gains reflect broader generation improvements or primarily a redistribution of probability toward preferred outputs. In this work, we systematically study eight representative preference-alignment methods from three complementary perspectives: preference improvement, broader generation capability, and responsible behavior. Using Best@ evaluation under increasing sampling budgets, we first confirm that alignment makes human-preferred outputs easier to sample. Beyond preference-oriented evaluation, however, these gains often diminish as the sampling budget increases, with the base model matching or surpassing aligned models on several benchmarks. Prompt-level analysis provides further insight: alignment does not uniformly improve success across prompts, but instead shifts the success distribution, increasing repeated-generation success for some prompts while leaving more prompts unsolved within the same budget. These changes also extend beyond capability, as most evaluated methods increase unsafe-generation rates and all increase occupational gender imbalance relative to the base model. Finally, we show that agentic generation with a frozen base model can achieve competitive generation performance while better preserving responsible behavior. Overall, our results suggest that preference alignment is better understood as reshaping the generation distribution toward preferred outputs, and motivate evaluating not only preference gains, but also how they transfer to broader capabilities and affect other model behaviors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.