acceptodds
Under review as a conference paper at ICLR 2027

DPMPC: Learning Diffusion Priors for Model Predictive Control

Abstract

Model predictive control (MPC) has achieved notable success in continuous control. MPC often uses a policy prior to propose candidate trajectories and to learn the value function. When the prior differs from the executed planner, this actor divergence biases the value estimates of candidate trajectories. The commonly used Gaussian prior cannot represent the non-Gaussian action distributions induced by stochastic planning, leaving a residual divergence. To address this, we propose DPMPC (Diffusion Priors for Model Predictive Control), which trains an expressive diffusion prior to match the planner's action distribution. Generalized lazy reanalyze estimates this distribution from multiple recent MPC realizations stored per replay state, without rerunning MPC at every policy update. An advantage-weighted score-matching objective then fits the prior to this estimate while emphasizing high-value actions. DPMPC achieves state-of-the-art asymptotic performance and sample efficiency across 14 continuous-control tasks from DMControl, MyoSuite, and HumanoidBench. Further analysis shows that DPMPC yields more accurate value estimates than Gaussian-prior baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.