Learning multi-step policy for planning via world-model-guided flow matching
Abstract
Planning with learned world models has demonstrated strong performance in online reinforcement learning for continuous control, but prominent TD-MPC2-style methods generate action sequences by repeatedly sampling from a one-step Gaussian policy conditioned on model-predicted states. This construction limits the policy's ability to capture diverse planner-generated actions, reduces temporal coherence across the full sequence, and makes sequence generation sensitive to model prediction errors. We propose FLOWP, which replaces the one-step Gaussian policy with a multi-step flow policy. We train the policy with world-model-guided flow matching, which leverages model-predicted returns to improve the policy while regularizing it toward the planner's action distribution. By jointly sampling entire action sequences, our method can represent diverse planner-induced distributions and temporal dependencies while reducing sensitivity to model prediction errors. We evaluate FLOWP on 7 high-dimensional DMControl tasks and 14 HumanoidBench tasks using the same hyperparameters across all tasks. FLOWP demonstrates strong sample efficiency, achieving the highest normalized task-average performance throughout most of training on both benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.