Spend Where It Matters: Particle Search with MH Refinement for Budget-Constrained Audio-Visual Alignment
Abstract
Reward-guided inference-time alignment can improve joint audio-visual generation, but optimizing multiple alignment objectives requires many expensive model evaluations, making existing methods costly under limited compute. Particle-based search enables global exploration but incurs estimation error with a finite population, while Metropolis-Hastings (MH) can refine candidates locally but benefits from a strong initialization. We introduce Spend Where It Matters (\modelname), a two-stage, training-free, reward-agnostic framework that combines particle search with MH refinement under a fixed compute budget. In the first stage, we adaptively allocate particles across denoising checkpoints using online variance estimates, minimizing estimation variance while concentrating compute where it is most useful. This search provides a promising candidate for MH refinement, which uses local noise-and-denoise proposals on completed outputs to further improve reward. By starting from a strong candidate, MH can recover additional reward without requiring a larger particle population or costly global exploration. We show that, at matched compute, combining particle search with MH refinement achieves higher rewards than either approach alone. On JavisBench-Mini, at closely matched compute, \modelname achieves mean relative gains across seven metrics of over BoN and over EvoSearch with JavisDiT++, and over BoN and over EvoSearch with LTX-2. Beyond joint audio-visual generation, we also demonstrate the benefits of adaptive particle allocation on SD1.5 image generation, improving AestheticScore by and ImageReward by over uniform allocation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.