Score-Based One-step MeanFlow Policy Optimization
Abstract
Diffusion and flow matching have emerged as expressive policy classes in reinforcement learning, but their reliance on multi-step denoising imposes substantial computational overhead at inference time, which is particularly problematic in online RL. MeanFlow offers a promising alternative by learning an average velocity field that maps noise to data in a single network evaluation. However, MeanFlow requires samples from the target distribution to construct its target velocity field, which are unavailable in online RL, and existing workarounds that use best-of- selections as training targets can only reinforce actions the current policy already produces. We propose Score-Based One-step MeanFlow Policy Optimization (SOM), an actor-critic algorithm that constructs the target velocity field directly from Q-function gradients via score estimation and the probability flow ODE, thereby transporting noise toward high-value modes. To stabilize score estimation under an evolving critic, we adapt the inverse temperature to control candidate-weight concentration via the effective sample size (ESS) across diffusion times, and anneal the target ESS over training. In fully online RL, SOM achieves the highest normalized return across five MuJoCo locomotion benchmarks with a single generation step, while substantially reducing inference latency compared to diffusion-based policies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.