acceptodds
Under review as a conference paper at ICLR 2027

Reward-Aligning Few-Step Flow Models with Integrated Regularizers

Abstract

Flow-based generative models are increasingly adapted to downstream rewards, preferences, and design objectives. Most existing approaches adapt the time-dependent drift of the underlying continuous flow. These continuous-time dynamics enable effective and probabilistically-principled reward alignment, but incur the cost of expensive simulations and require the challenging propagation of terminal reward gradients backwards through the trajectories. Recent accelerated one- and few-step samplers raise the prospect of improved fine-tuning efficiency, however such approaches side-step the continuous dynamics that are useful for effective and stable regularization. In this work, we propose a method for fine-tuning a flow map that preserves the advantages of both regimes, while also retaining the ability to perform any-step generation. Our approach directly optimizes the few-step generations using gradients of the reward function, while regularizing using a family of integrated divergences along the continuous-time probability flow induced by the model. Experiments on image and protein generation show that the proposed approach combines the efficiency of few-step generation with the stability of continuous-time regularization along the probability flow path.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.