Explore Before You Optimize: Robust Initialization for Reward-Guided Generation
Abstract
Reward-guided test-time optimization improves generation quality but often reduces diversity by concentrating samples around a limited set of high-reward modes. Recent explore-then-refine methods encourage early exploration before reward-guided refinement, yet this process can be inefficient and the resulting diversity may still degrade during subsequent optimization. We address these limitations by constructing a robust, coverage-oriented initialization population. Building on guidance potential, we introduce a local robustness objective that favors latents located in broad low-potential neighborhoods, where the low-contraction property is more stable under reward-driven perturbations. We further propose a population-level semantic coverage objective that reduces redundancy in the predicted representation space, allowing a limited candidate set to span a broader range of semantic modes. Experiments across multiple benchmarks show that our method achieves competitive generation quality while consistently improving output diversity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.