Source Portfolio Guidance for Reward-Aligned Diffusion Generation
Abstract
Test-time guidance for text-to-image synthesis seeks to align a pretrained generator with task-specific preferences while retaining diverse outputs, without updating model parameters. Gradient-free source-space methods enable optimization with black-box rewards, but their active populations are designed to continue search rather than to represent all useful discoveries. Consequently, reward-ranked output selection may discard distinct realizations encountered along the search trajectory and fill a small returned set with similar high-scoring samples. We propose Source Portfolio Guidance (SPG), which explicitly separates candidate discovery from finite-set construction under a reward constraint. Trajectory Support Consolidation (TSC) combines multiscale reward-driven proposals with a persistent history, preserving evaluated alternatives beyond the compact parent archive used for exploration. Anchored Support Compression (ASC) retains the best observed candidate and applies winner-anchored farthest-first coverage to a reward-qualified pool, selecting complementary representatives without additional generator evaluations. For fixed evaluated records, this construction preserves peak reward and provides a classical coverage bound when the eligible pool can fill the output set. Across Stable Diffusion 1.4, 1.5, and XL, SPG improves best-of-four ImageReward by 0.105–0.176 and CLIP diversity by 25.2–27.2% over previous state-of-the-art method Source Parallel Tempering (SPT) under the same budget of 420 candidates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.