IQ-Diff: Iterative Q-guided Sampling-Time Diffusion Alignment with Black-Box Physics Rewards
Abstract
Diffusion models provide generative priors for physical prediction, but statistically plausible outputs need not satisfy downstream physical objectives. We study sampling-time alignment when feedback is available as a scalar terminal reward from a black-box physics simulator. IQ-Diff keeps the diffusion backbone frozen and repeatedly learns action-value functions from trajectories of the current guided sampler. Re-evaluation accounts for changes in the continuation policy, while accumulated Q-scores implement successive KL-regularized updates over local denoising candidates. Deployment uses these learned scores without simulator gradients or further simulator calls. We analyze the finite-candidate selection process directly, establishing a final-iterate bound under controlled Q-estimation error and separating this optimization error from the proposal's coverage limitation. Exact current-policy evaluation yields monotonic reward improvement and an final-iterate guarantee relative to the best selector with the same candidate budget, where is the number of outer iterations. Experiments on fluid dynamics and thermal-process prediction demonstrate improved prediction accuracy and favorable quality–cost tradeoffs against unguided sampling, final-sample search, physics-reference guidance, single-Q guidance, and SMC-style resampling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.