Simple Flow Gradient: Fine-Tuning Flow Policies with the Direct Q-Gradient
Abstract
Generative policies trained by imitation rarely reach the reliability that deployment demands, and reinforcement learning (RL) offers a way to improve them through feedback from their own interactions. Off-policy methods reuse past experience for sample efficiency, but applying them to a generative policy raises two concerns identified by prior work: backpropagating the critic's gradient through the sampling chain may be unstable, and an update without an anchor to the data may exploit errors in the critic's value estimates. This work examines whether the direct Q-gradient, the most elementary such update, can fine-tune a pretrained generative policy without an anchor. Analyzing direct Q-gradient updates on a one-step flow policy, we find that a single sufficiently large update can erase pretrained competence. This motivates bounding the action space change of each episode's update while allowing continued adaptation beyond the pretrained policy. Simple Flow Gradient (SFG) applies this bound by projection to a pretrained flow policy updated with the direct Q-gradient; the bound is its only addition to the actor-critic recipe, which keeps no anchor term, distillation stage, or auxiliary network. Starting from converged behavior cloning policies under matched interaction budgets, SFG reaches a mean task completion score of 0.98 on five manipulation tasks from Robomimic and Franka-Kitchen, and a score of 0.8 more than twice as fast as the fastest baseline. Additional experiments show that SFG also applies to multistep flow policies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.