Policy Optimization via Drifting
Abstract
Reinforcement learning (RL) of diffusion and flow-based policies has become a popular paradigm in robot learning, where actions are continuous and multimodal. Sampling from these generative policies, however, costs many sequential network evaluations and incurs latency, making them a poor fit for dynamic tasks with tight feedback loops. Drifting models offer an appealing alternative. Empirically, they are as expressive as diffusion and flow models, but they can produce a sample with a single network evaluation. In this work, we propose *reward drifting*, an on-policy policy gradient optimization method. Starting from the theory of Wasserstein gradient flows, we derive a single fixed-point objective and justify that optimizing it is equivalent to increasing the expected discounted reward. Our method does not require access to action-likelihoods or surrogates. Empirically, we evaluate our algorithm across a range of manipulation tasks, including Push-T, dexterous in-hand reorientation, and long horizon bimanual manipulation. Our method offers substantially faster inference than diffusion and flow-matching baselines fine-tuned with existing RL methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.