PLATO: A Unified World-State Representation for VLA Perception and Prediction
Abstract
As in Plato’s allegory of the cave, vision-language-action (VLA) models learn from 2D visual projections, while the underlying physical world state remains implicit. We introduce PLATO (Physical Language for Action through Tokenized Observations), a unified world-state representation for VLA perception and prediction. Learned from multi-view observations with geometric, semantic, and forward–inverse dynamics supervision, PLATO preserves 3D structure, semantics, and interaction-relevant information in a compact form. Continuous features or embeddings of discrete codes provide current-state context, and 128 discrete state tokens serve as autoregressive future-state targets alongside action supervision. Both roles leverage the VLM’s native conditioning and next-token prediction paradigm, without specialized dense prediction branches. Across VLA pretraining and post-training on three simulation benchmarks and four real-world tasks, PLATO improves manipulation success and robustness. Future-state supervision raises average real-world success by 11.7 and 25.4% under nominal and randomized conditions, respectively, and current-state context further improves viewpoint robustness, with a 40% gain over supervision alone in flower insertion. Beyond performance, our analyses show that PLATO reshapes VLA representations: geometry becomes more linearly accessible and RGB embeddings more consistent across viewpoints, revealing that explicit world-state inputs and targets substantially alter the physical structure learned by the VLA from robot data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.