Do Semantic Feature Spaces Make Better World Action Models?
Abstract
World action models (WAMs) offer a promising approach to robotic control by coupling future visual prediction with action generation. Many existing WAMs predict future states in reconstruction-oriented variational autoencoder (VAE) la- tent spaces, raising the question: do semantic feature spaces provide a better foun- dation for robot control? We investigate this question through a systematic com- parison of the Stable Diffusion 3.5 VAE, DINOv2, DINOv3, and SigLIP 2 within a common cascaded WAM framework. We use flow matching to predict future features directly in each encoder’s representation space and an inverse dynamics model to decode actions from these predictions, with training configurations tuned for each representation. We analyze action decoding, future dynamics prediction, and closed-loop task success across training data and compute budgets. Experi- ments on RoboCasa365 and transfer from real-world DROID data to high-fidelity simulation in RoboLab assess WAMs’ performance and generalization. Across the evaluated setups, semantic-space WAMs achieve higher task success and scale more favorably with compute, data, and pretraining on additional tasks than VAE- based WAMs. These findings support semantic feature spaces as a stronger foun- dation for effective and scalable world action modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.