Controlling Image-Space Drift in Inference-Time Noise Optimization
Abstract
Inference-time alignment steers pre-trained diffusion models toward target rewards without updating model weights. Among these methods, optimizing the input noise is especially powerful: by directly modifying the generation process, it can reach rare, high-reward samples. However, this flexibility also makes reward hacking significantly worse, often producing high-scoring images that lose the structure and realism of the base model. Existing approaches try to prevent this by constraining the optimized noise, but we show this is insufficient: the denoising process can amplify small changes in noise into large differences in the generated images. Thus, we propose BIDS (Bounding Image-space Divergence via Score functions), which uses a score-based constraint to limit image-space drift during noise optimization. BIDS uses the Data Processing Inequality to relate image-space divergence to divergence over denoising trajectories, then computes a practical surrogate using only the frozen model's score function. This preserves image structure during reward optimization. Across UNet and DiT architectures, BIDS reduces distributional drift (CMMD) by over 75% and achieves up to an 80.30% human preference win rate. We further extend BIDS to pairwise preferences, enabling alignment to non-differentiable human feedback.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.