RACE: Reward-Guided Adaptive Candidate Exploration for Sample-Efficient Diffusion Fine-Tuning
Abstract
Reward-guided diffusion fine-tuning requires informative candidates under a limited scoring budget. Independent sampling provides coverage, while a shared scalar perturbation cannot adapt exploration across latent coordinates. We introduce Reward-Guided Adaptive Candidate Exploration (RACE), which learns a state-conditioned, entry-wise perturbation map to construct additional candidates around selected generated trajectories while retaining independent base samples. When paired training images are available, RACE uses deterministic inversion to provide reference-derived exploration centers within the same candidate budget. Both variants jointly optimize the generator and perturbation predictor through a single reward-weighted objective. Our analysis derives sufficient conditions for improved expected candidate reward, separating parent-selection benefits and perturbation effects and quantifying the roles of reference quality and inversion error. Experiments cover natural-image generation, including a multi-topic COCO subset, remote sensing, and medical imaging. Both variants achieve higher mean scores than DiffusionNFT across all five COCO evaluators. On brain MRI, RACE achieves a joint score of , compared with for DiffusionNFT. These findings support learning candidate construction as a means of improving diffusion fine-tuning under matched scoring budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.