Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
Abstract
Online off-policy reinforcement learning (RL) is governed by two coupled design dimensions: policy representation (e.g. Gaussian vs generative policies) and optimisation methodology (e.g. soft actor-critic (SAC) vs mirror descent (MD)). Gaussian policies offer fast inference and tractable entropy estimation but have limited ability to model multimodal action distributions, whereas generative policies provide richer action distributions at the cost of iterative sampling or intractable entropy estimation. On the optimisation side, SAC-style entropic exploration and MD updates can be interpreted as minimising distinct Kullback-Leibler divergences: SAC-style exploration performs soft policy improvement towards a value-induced Boltzmann distribution, whereas MD constrains updates to remain close to the previous policy. Combining SAC-style entropy and MD regularisation supports exploration and conservative policy updates, but can yield multimodal targets beyond the representational capacity of unimodal Gaussian policies. This mismatch calls for an expressive policy with efficient sampling and explicit entropy control. To this end, we introduce Stochastic MeanFlow Policies (SMFP), a one-step generative policy class that combines MeanFlow-based noise-to-action mappings with Gaussian reparameterisation. The latent mapping represents diverse action modes, while the learnable conditional Gaussian scale yields an analytic entropy lower bound. We use this bound to regulate stochasticity alongside direct value guidance and advantage-weighted MeanFlow regression, making entropy-regularised mirror-descent learning practical with one-step generative policies. Empirically, on seven MuJoCo benchmarks, SMFP achieves strong performance against both Gaussian and generative baselines while maintaining single-step inference efficiency. Complementary experiments demonstrate multimodal action modelling and effective exploration in sparse-reward navigation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.