acceptodds
Under review as a conference paper at ICLR 2027

Flow-Time Bellman Propagation for Generative Policy Improvement

Abstract

Flow policies represent complex action distributions by transforming noise into actions through a sequence of generation steps. To learn useful behavior, these policies must incorporate value information into the generation process. Value-weighted flow matching favors high-value training actions, but leaves the target direction for each noise–action pair unchanged. We propose Flow-Time Bellman Propagation (FTBP), which learns value feedback for intermediate generation states and uses it to train the flow. A continuation potential predicts the critic value of the action obtained by completing the current flow. We train this potential with endpoint critic supervision and local bootstrap targets, avoiding a separate completed rollout for each intermediate target. The learned potential then guides training through successor evaluation (FTBP-S), target shaping (FTBP-T), or their combination (FTBP-C). During online fine-tuning, recorded generation traces provide additional supervision under the same policy objectives. Our analysis relates the two updates and characterizes value propagation for a matched finite-grid sampler. Experiments demonstrate that FTBP delivers strong and consistent improvements in both offline learning and online fine-tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.