Rethinking the Representation Space for World-Action Models: Pixels, VAE Latents, or Semantic Representations?
Abstract
World-action models (WAMs) built on video generation models learn interaction dynamics that map observations to robot actions. Because video generators can produce implausible or hallucinated futures under perturbation, many recent WAMs add auxiliary supervision on other representations to improve robustness under distribution shift. These additions increase post-training cost and introduce trade-offs among loss weights, while leaving open where generalization comes from. We instead ask a more basic question: in which representation space should a WAM predict the future? Within a shared framework and without auxiliary objectives, we compare models that predict the future in pixels (Pixel-WAM), VAE latents (VAE-WAM), or semantic features from frozen pretrained encoders (RAE-WAM). Across six fusion designs, hidden-state fusion transfers video pretraining most effectively. Following this, the V-JEPA RAE-WAM reaches 55.2% success from scratch, compared with 45.7% for the DINOv2 RAE-WAM, 45.4% for VAE-WAM, and 37.2% for Pixel-WAM on LIBERO-Plus. The same representation ordering holds in clean simulated bimanual manipulation, and on a real dual-arm platform the V-JEPA RAE-WAM outperforms π0.5 and DreamZero. Our results point to a simple principle for building world-action models: rather than reconstructing pixels or relying on auxiliary supervision, predict the future in a pretrained semantic space whose structure already reflects how scenes evolve under action.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.