Diversity Preference Ranker: Diverse Text-to-Image Generation with Unified Scoring
Abstract
We explore diverse image generation from accelerated text-to-image models, whose outputs can remain similar across different random seeds. Existing selection methods typically combine independently trained diversity and quality scorers that encode each candidate in isolation. We argue that diversity and preference scoring can benefit from a shared visual representation, allowing both supervision signals to shape the features used for scoring. We further emphasize that capturing a candidate's distinctiveness requires considering the other images in the pool. These considerations motivate Diversity Preference Ranker (DPR), a lightweight ranker that replaces separate diversity and preference scorers with a shared representation of the candidate pool. Self-attention encodes all candidates together into a shared field that provides pool-conditioned pairwise diversity distances and per-candidate preference scores. Spatial feature distillation and human preference comparisons supervise the shared representation. Supervising diversity and preference individually does not directly teach the ranker how to combine them when selecting candidates. We therefore introduce listwise selection distillation, which trains the combined scores to imitate a teacher's next-candidate choices given the selected subset. Extensive experiments across four accelerated generators show that the same frozen ranker achieves comparable diversity and preference to the evaluated selection baseline without generator-specific retraining. As a further application, we demonstrate that the same frozen ranker also supports noise optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.