World Tokens: Enhancing Embodied Policies with Training-Time World Modeling
Abstract
Vision-language-action (VLA) models are efficient at inference but do not explicitly model how physical scenes evolve as a task unfolds, whereas world-action models (WAMs) capture this evolution with pretrained video world models at the cost of keeping future generation or a large video backbone in the control loop. We introduce World Tokens, an embodied policy architecture that uses world modeling only during training, through an exclusive predictive interface. A World Adapter transforms VLM features into a fixed set of world tokens, which condition a jointly fine-tuned future-video denoiser and serve as the action expert's sole visual-language context. The denoiser receives only a coarse layout cue of the current scene, so neither it nor the action expert can bypass the world tokens, and future-video supervision directly shapes the representation used for control. At deployment, the world-model branch is removed, leaving only the VLM, World Adapter, and action expert. We evaluate World Tokens on LIBERO, the zero-shot perturbations of LIBERO-Plus, SIMPLER, and real-world manipulation with a Galaxea R1 Pro robot. With a 2B backbone and no embodied action pretraining, it performs strongly across these benchmarks and consistently improves over a matched action-only baseline, while running at VLA-level latency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.