acceptodds
Under review as a conference paper at ICLR 2027

Drift, Then Spread: Learning One-Step Generative Policies for Online Reinforcement Learning

Abstract

Generative policies offer expressive action distributions for online reinforcement learning (RL), but many diffusion- and flow-based policies require costly iterative sampling. We propose DriftSpread, a one-step generative policy based on Drifting for online maximum-entropy (MaxEnt) RL. To incorporate entropy without a tractable policy density, we build on an energy-based view of Drifting and derive critic-guided Drift and entropy-induced Spread from the negative MaxEnt objective. Our method learns these updates sequentially in the same actor, using action regression for value improvement and distribution matching to capture the effect of entropy. The actor updates require neither direct samples from the critic-defined Boltzmann target nor policy-density or score estimation. Under exact fitting with a fixed critic, we prove convergence to the Boltzmann target in the small-step limit followed by the long-time limit. On MuJoCo, DriftSpread achieves the highest aggregate normalized return among all evaluated methods, while outperforming the diffusion- and flow-based baselines in both training and inference efficiency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.