Predicting Semantic Latents for Robust and Efficient World-Action Models
Abstract
World-Action Models (WAMs) adapt pretrained video generative models for robot control by jointly predicting future pixel observations and actions. We argue that robot manipulation requires reasoning about high-level scene structure and object dynamics, but does not require pixel-level fidelity, which can be sensitive to nuisance factors like lighting. Rather than predicting future pixels and actions, we train our model to predict future semantic latents (e.g., frozen DINO, SigLIP, or V-JEPA encodings of future frames) and future actions. We refer to our approach as Semantic Latent Action Prediction (SLAP). Importantly, we find that video diffusion backbones can be efficiently post-trained to directly regress (in a single forward pass) future semantic latents, eliminating the need for expensive iterative video denoising. Semantic targets improves robustness to environmental perturbations (like lighting or even end-effector initial position) and requires far less robot data for post-training. We ablate several design choices and find that intermediate layers of the pretrained video transformer provide better features for control compared to the final layers, motivating an early-readout design that further improves efficiency and performance. On standard simulation benchmarks like LIBERO and RoboTwin, SLAP remains competitive with strong VLA and WAM baselines that train on 50× more robot-data. On the perturbation-heavy LIBERO-Plus benchmark, SLAP substantially outperforms pixel-level prediction. Finally, we show that SLAP can be deployed for real-world robot manipulation with as little as 30 demonstrations. SLAP is 1.9× faster than Fast-WAM, 2.2× faster than Flex-π Joint, and 19× faster than LingBot-VA, while remaining comparable in speed to .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.