acceptodds
Under review as a conference paper at ICLR 2027

Explore Before You Optimize: Robust Initialization for Reward-Guided Generation

Abstract

Reward-guided test-time optimization improves generation quality but often reduces diversity by concentrating samples around a limited set of high-reward modes. Recent explore-then-refine methods encourage early exploration before reward-guided refinement, yet this process can be inefficient and the resulting diversity may still degrade during subsequent optimization. We address these limitations by constructing a robust, coverage-oriented initialization population. Building on guidance potential, we introduce a local robustness objective that favors latents located in broad low-potential neighborhoods, where the low-contraction property is more stable under reward-driven perturbations. We further propose a population-level semantic coverage objective that reduces redundancy in the predicted representation space, allowing a limited candidate set to span a broader range of semantic modes. Experiments across multiple benchmarks show that our method achieves competitive generation quality while consistently improving output diversity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.