On the Representation of World Action Models
Abstract
World action models (WAMs) use visual prediction to learn intermediate representations (e.g., global or dense semantics, geometry, and spatiotemporal dynamics) for robot action prediction. Learning these representations depends on which visual features are used as inputs/outputs and how these features are predicted. In this paper, we conduct comprehensive analyses on how these two choices shape intermediate representations and thus help action prediction. Our study yields two key findings:1) V-JEPA 2.1 features (which are good at spatiotemporal dynamics) outperform Wan VAE (pixel reconstruction), DINOv3 (dense semantics), and SigLIP-2 (global semantics) features for robot action prediction. 2) using frame-wise autoregression for visual feature prediction performs better than diffusion, i.e., shaping the intermediate representations better for action prediction. On RoboTwin 2.0, our representation recipe (V-JEPA 2.1 + frame-wise autoregression) achieves an average success rate of 85.9%, compared with 38.0% for the baseline recipe (Wan VAE + diffusion), without additional robot pretraining. This advantage persists after additional pre-training on 3,300 hours robot data, with success rates of 91.5% (our recipe) and 75.1% (baseline recipe). We hope these findings provide helpful guidances for developing more effective representation recipes for WAMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.