acceptodds
Under review as a conference paper at ICLR 2027

On the Representation of World Action Models

Abstract

World action models (WAMs) use visual prediction to learn intermediate representations (e.g., global or dense semantics, geometry, and spatiotemporal dynamics) for robot action prediction. Learning these representations depends on which visual features are used as inputs/outputs and how these features are predicted. In this paper, we conduct comprehensive analyses on how these two choices shape intermediate representations and thus help action prediction. Our study yields two key findings:1) V-JEPA 2.1 features (which are good at spatiotemporal dynamics) outperform Wan VAE (pixel reconstruction), DINOv3 (dense semantics), and SigLIP-2 (global semantics) features for robot action prediction. 2) using frame-wise autoregression for visual feature prediction performs better than diffusion, i.e., shaping the intermediate representations better for action prediction. On RoboTwin 2.0, our representation recipe (V-JEPA 2.1 + frame-wise autoregression) achieves an average success rate of 85.9%, compared with 38.0% for the baseline recipe (Wan VAE + diffusion), without additional robot pretraining. This advantage persists after additional pre-training on 3,300 hours robot data, with success rates of 91.5% (our recipe) and 75.1% (baseline recipe). We hope these findings provide helpful guidances for developing more effective representation recipes for WAMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.