Flow Matching Loss for Offline Flow Policy Reinforcement Learning
Abstract
Flow models have demonstrated substantial promise for robotics control tasks. Usually, they are trained using behavior cloning, and a key challenge in training these policies is adapting traditional offline RL algorithms to flow policies. In particular, modern offline RL algorithms include a deterministic Q-gradient, which requires backpropagating through the iterative decoding process. This step undermines one of the key advantages of flow policies, which is that every decoding step is trained using a strong supervised signal derived directly from the final output. We propose the following solution: rather than target the action in the offline dataset, we first use Q-guided inference-time policy optimization to identify a better action, and then directly train the flow policy to target that action using the standard flow matching objective. We implement this idea in the context of Reversal Q-learning (RQL), a state-of-the-art offline RL algorithm for flow policies, to obtain Q-gradient RQL (QGRQL). Across multiple tasks from the OGBench locomotion and manipulation environments, we show that QGRQL outperforms existing state-of-the-art methods. Our results demonstrate the promise of better aligning offline RL algorithms with flow matching objectives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.