acceptodds
Under review as a conference paper at ICLR 2027

Model-based Flow Policy Optimization

Abstract

Generative control policies (GCPs), such as diffusion and flow matching policies can represent expressive, multi-modal action distributions, but training them with reinforcement learning is difficult: their likelihoods are expensive to compute, and backpropagating through the sampling chain is costly and unstable. Existing methods sidestep this with model-free, on-policy RL, which is prohibitively sample-inefficient, especially from pixels. We observe that a learned world model resolves this tension. Likelihood-free objectives such as advantage-weighted flow matching require abundant on-policy data, which a latent world model can generate cheaply in imagination. We propose MB-FPO, which trains flow policies by regressing onto imagined actions weighted by their advantages, with no likelihoods and no backpropagation through sampling. With temporally correlated exploration noise, MB-FPO is substantially more sample efficient than prior flow policy optimization methods on the DeepMind Control and Meta-World benchmarks, outperforms model-free baselines on most pixel-based tasks, and matches or exceeds a model-based Gaussian policy trained on the same world.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.