acceptodds
Under review as a conference paper at ICLR 2027

STEP: Post-Training Few-Step Flow Models by Propagating Reward Gradients Through Every Generation Step

Abstract

We study reward fine-tuning of few-step Gaussian-mixture (GM) flow policies. Each network evaluation defines a velocity field over a sampling interval, allowing analytic estimates of the clean latent at intermediate states. We introduce STEP, which uses these estimates as probes for constructing surrogate reward gradients. One probe is sampled within each deployment interval, and a straight-through chain retains dependencies between intervals. Weighted fusion connects all probes to a shared terminal reward while preserving the sampled image in the forward pass. The probes require no additional transformer evaluations. The full training objective combines this terminal loss with auxiliary rewards and frozen-teacher velocity regularization. After optimization steps on FLUX.1-dev, STEP achieves GenEval, OCR, and held-out HPSv3 with four neural function evaluations. The official distilled student scores , , and , respectively. Within the full objective, removing the chain or replacing interior probes with zero-depth estimates reduces OCR and GenEval. These results support using the interval structure of GM policies for reward fine-tuning at a fixed four-step deployment budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.