Improving Truncated First-Order Gradients for Flow Policies with Surrogate Objective
Abstract
Differentiable simulation offers a way to improve the sample efficiency of reinforcement learning for robotics. However, correctly utlizing the first-order information from the simulator for training vision-based flow-matching policies remains challenging. Visual observations may be nondifferentiable with respect to the environment state, and iterative action generation in flow matching makes end-to-end backpropagation costly. Truncating the observation gradient path makes the analytical policy gradient (APG) tractable, but the resulting truncated APG is biased. Likelihood-ratio policy gradients (LRPGs) avoid this differentiable requirement, but can suffer from high variance. We introduce Differentiable Flow Policy Optimization (DFPO), a method that combines these complementary signals to improve policy updates under incomplete differentiation. DFPO uses LRPG, estimated with a flow-matching surrogate objective, to improve the update direction of the truncated APG. On classical control problems, we show that the combined update has low variance and is better aligned with the true policy gradient. We further evaluate DFPO on robotic tasks with visual and state observations. In the vision-based setting, DFPO consistently outperforms baselines using either truncated APG or LRPG alone. In the state-based setting, it achieves performance comparable to methods using the full APG. These results demonstrate that DFPO, which combines the truncated APG with the LRPG can support effective reinforcement learning of vision-based flow-matching policies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.