acceptodds
Under review as a conference paper at ICLR 2027

ROOT: Resolving Coefficient Ambiguity in Endpoint Reward Tuning for Flow Models

Abstract

Forward-process reward tuning uses rewards on behavior-generated endpoints to define a reward-improved endpoint distribution, then re-noises these endpoints to learn its induced conditional flow. We identify coefficient ambiguity in stochastic velocity training, where positive weights defining the same endpoint distribution and induced flow can produce different updates. Condition-wise reward centering exposes this ambiguity by rescaling exponential weights without changing the target at a fixed tilt. We show that the target and behavior distributions uniquely determine a mean-one target-to-behavior density ratio. ROOT uses this probability-defined ratio as the coefficient for stochastic velocity learning, making its construction invariant to condition-wise rescaling of endpoint-equivalent weights. Its single-pass component, RootSolve, yields a behavior-centered finite-tilt objective whose conditional population optimum matches the target flow under behavior-flow consistency. Its multi-pass component, RootTrace, reuses scored rollouts through moving-reference updates that retain reward-dependent corrections on later passes while removing the pull toward the original behavior prediction. On Stable Diffusion 3.5 Medium, ROOT reaches matched DiffusionNFT training-reward thresholds with – and – fewer outer iterations across single- and multi-pass settings, respectively, and achieves comparable final performance under sequential multi-reward training with about % fewer iterations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.