Hamiltonian World Model for Physically Consistent Embodied Prediction
Abstract
Embodied world models are increasingly expected to serve not only as video generators, but as internal simulators for action-conditioned prediction. However, most existing models parameterize latent transitions with generative modules, leaving physical structure to be learned implicitly from data. This design may accumulate physical drift over long horizons, especially in contact-rich manipulation. We propose Hamiltonian World Model (), a physically structured latent transition framework. The core concept of is to move physical inductive bias from auxiliary regularization or post-hoc evaluation into the transition law itself: backbone latents are mapped into a learned phase space of generalized coordinates and momenta, evolved through dissipative Hamiltonian dynamics. A bidirectional phase-space interface and a multi-step transition-consistency objective keep the structured rollout compatible with existing world-model encoders and decoders, making a drop-in replacement for the latent transition module. In simulation, improves speed, acceleration, and jerk prediction while better preserving system energy than unstructured transitions. We further instantiate across three embodied world-model settings: Ctrl-World for single-arm multi-view manipulation, EVAC for dual-arm single-view manipulation, and Cosmos-Predict2.5-2B for single-arm single-view task execution. On Robocasa, Cosmos-Predict2.5-2B with achieves relative success-rate gains of 2.7%; on real-robot AgiBot World data, delivers relative gains of 9.7% on nDTW with Ctrl-World and 8.4% on Logic with EVAC. These results suggest that physically structuring the latent transition is a practical path toward embodied world models whose imagined futures are more stable, more physically faithful, and more useful for action. The project page is available at https://hamilton-world-model.github.io/Hamiltonian_World_Model/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.