acceptodds
Under review as a conference paper at ICLR 2027

What Matters for Value Guidance of Flow Policies in Reinforcement Learning?

Abstract

Flow policies are increasingly used to model behavior policies in offline reinforcement learning, but their iterative generative process complicates value-based policy improvement. Value guidance offers an alternative that leaves the behavior policy parameters unchanged by steering the generation process with the gradient of a learned critic. Existing value guidance methods instantiate this idea in diverse ways, differing in the flow timestep at which guidance is applied, how many guidance steps are taken, and how the guidance signal is constructed from the critic. Because these factors are rarely varied or studied in isolation, it is unclear which of them drive performance. In this work, we first identify which of these factors matter through a controlled study that varies each independently, and then examine how the findings change when the critic is trained on actions produced by the guidance itself. We find that the guidance signal determines which flow timesteps admit effective guidance, while the benefit of additional guidance steps differs across tasks. A single guidance step often suffices in tasks where the behavior cloning (BC) policy already performs well, whereas more steps are needed to substantially improve performance in tasks where the BC policy rarely succeeds. We also find that when the critic learns from bootstrap actions generated by the guidance itself, the benefit of additional guidance steps saturates earlier, with only a few steps recovering most of the attainable improvement. Building on this, we show that applying only a few guidance steps at selected flow timesteps matches or exceeds guiding at every flow step while reducing inference cost by up to 63%, showing that these insights translate into practical and efficient value guidance. These findings suggest that guidance schedules and value learning should be designed together to exploit the iterative structure of flow policies for effective value guidance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.