Back to State: Explicit Physical State World Models for Continuous Control
Abstract
Temporal difference learning for model predictive control (TD-MPC) and its recent variants achieve strong continuous control performance by learning task-oriented latent dynamics and performing short-horizon planning entirely in latent space. For state-based control, however, this removes the physical semantics from imagined trajectories and makes physical state predictions unavailable for deployment-time objectives. Prior work has attempted to recover explicit state predictions from latent world models, but reconstruction-based approaches can suffer substantial performance degradation on challenging high-dimensional continuous control tasks. In this work, we introduce State TD-MPC (S-TD-MPC), which directly propagates the normalized physical state while augmenting it with learned latent features recomputed at each imagined step, providing explicit physical state predictions while retaining strong representation capacity. We further align value learning with MPC through MPC-based Bellman targets and introduce reanalysis via Bellman target bounds to refresh stale planner targets selectively. Beyond the training objective, we study three capabilities enabled by explicit state predictions: zero-shot transfer across tasks with shared dynamics, deployment-time state constraints, and analytic reward evaluation during training. Experiments on challenging continuous control tasks show that S-TD-MPC matches the performance of state-of-the-art methods in the TD-MPC family, achieving an average return 2.2% higher than the strongest baseline, while enabling these constraint-aware and transferable control capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.