acceptodds
Under review as a conference paper at ICLR 2027

Bridging Candidate Selection and Mean-Shift Guidance for Diffusion Reward Alignment via Linear Noise Ranking

Abstract

Training-free reward alignment of diffusion models follows two routes. Candidate selection draws candidate next states at each denoising step and keeps the one with the highest estimated reward, spending neural evaluations per step. Mean-shift guidance instead moves the transition mean along a single reward gradient. We bridge these two regimes by utilizing a linear functional towards a selected direction as a proxy reward model to rank the Gaussian noises at each steps and select according to those scores. Scoring each noises only requires one inner product, bypassing the neural network forward costs. Our analysis locates this rule precisely between the two routes. We propose a single representation that unifies the candidate selection method and the mean shift method, and view their difference only by the distribution of the strength scalar towards the reward optimizing direction. This enables us to derive exact leading order predictions of the relative reward optimization gaps between candidate selection method and mean shift method, which are validated through measurements. We further provide practical methods to compute the reward optimizing direction: a corrected direction whose benefit it predicts, and a least-squares slope, fitted from the same reward queries, that extends the rule to non-differentiable rewards. Our experiments show that our linear noise ranking matches or exceeds per-candidate selection at equal or lower neural cost and ties with strength-matched mean shift method as predicted.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.