Anisotropic Stochasticity for Inference-Time Scaling of Flow Models
Abstract
Inference-time reward alignment explores multiple candidates during generation and favors those with higher intermediate rewards. Since flow models generate samples deterministically, such exploration relies on marginal-preserving stochastic differential equations (SDEs), which let duplicated candidates branch from a shared intermediate state into distinct trajectories. Prior works add Gaussian noise with equal variance across all dimensions, but much of this isotropic noise falls on directions that barely change the final generated sample. We analyze how such perturbations translate into variation among the predicted endpoints, explaining why isotropic noise can lead to inefficient exploration. Building on this analysis, we propose an anisotropic stochastic transition framework formulated as an SDE that preserves the original flow marginals. Our framework shapes the noise with the posterior covariance of the current state, allocating more noise to directions where perturbations change the final sample more. We further show that the drift adjustment required to preserve the original marginals reduces to an acceleration of the flow. Experiments on text-to-image reward alignment and offline-RL demonstrate improved reward–compute trade-offs, even after accounting for the additional computational overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.