Model-based Flow Policy Optimization
Abstract
Generative control policies (GCPs), such as diffusion and flow matching policies can represent expressive, multi-modal action distributions, but training them with reinforcement learning is difficult: their likelihoods are expensive to compute, and backpropagating through the sampling chain is costly and unstable. Existing methods sidestep this with model-free, on-policy RL, which is prohibitively sample-inefficient, especially from pixels. We observe that a learned world model resolves this tension. Likelihood-free objectives such as advantage-weighted flow matching require abundant on-policy data, which a latent world model can generate cheaply in imagination. We propose MB-FPO, which trains flow policies by regressing onto imagined actions weighted by their advantages, with no likelihoods and no backpropagation through sampling. With temporally correlated exploration noise, MB-FPO is substantially more sample efficient than prior flow policy optimization methods on the DeepMind Control and Meta-World benchmarks, outperforms model-free baselines on most pixel-based tasks, and matches or exceeds a model-based Gaussian policy trained on the same world.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.