Forward Dynamics Models
Abstract
Vision-language-action (VLA) models and world action models (WAMs) have advanced robotic policy learning by building on pretrained vision-language models (VLMs) and video generative models (VGMs), respectively. Yet VLAs learn without explicit supervision from future world states, while WAMs that condition actions on predicted futures incur extra inference cost and let visual prediction errors propagate into actions. We present Forward Dynamics Models (FDM), a family of robotic action models that jointly predicts actions and future world states through an action-to-future paradigm. Future prediction is conditioned on action representations, while action prediction remains independent of predicted futures. End-to-end training lets ground-truth future states supervise the policy through these action representations. At inference, FDM skips future generation and runs as efficiently as a standard VLA. The paradigm is backbone-agnostic, and we instantiate it with both VLM and VGM backbones. Experiments in simulation and on real robots demonstrate the effectiveness of FDM across both backbone families, achieving 99.4% average success on LIBERO and gains of up to 13 percentage points over VLA and WAM methods on real-world tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.