acceptodds
Under review as a conference paper at ICLR 2027

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

Abstract

Vision-language-action (VLA) models are efficient at inference but do not explicitly model how physical scenes evolve as a task unfolds, whereas world-action models (WAMs) capture this evolution with pretrained video world models at the cost of keeping future generation or a large video backbone in the control loop. We introduce World Tokens, an embodied policy architecture that uses world modeling only during training, through an exclusive predictive interface. A World Adapter transforms VLM features into a fixed set of world tokens, which condition a jointly fine-tuned future-video denoiser and serve as the action expert's sole visual-language context. The denoiser receives only a coarse layout cue of the current scene, so neither it nor the action expert can bypass the world tokens, and future-video supervision directly shapes the representation used for control. At deployment, the world-model branch is removed, leaving only the VLM, World Adapter, and action expert. We evaluate World Tokens on LIBERO, the zero-shot perturbations of LIBERO-Plus, SIMPLER, and real-world manipulation with a Galaxea R1 Pro robot. With a 2B backbone and no embodied action pretraining, it performs strongly across these benchmarks and consistently improves over a matched action-only baseline, while running at VLA-level latency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.