acceptodds
Under review as a conference paper at ICLR 2027

PLATO: A Unified World-State Representation for VLA Perception and Prediction

Abstract

As in Plato’s allegory of the cave, vision-language-action (VLA) models learn from 2D visual projections, while the underlying physical world state remains implicit. We introduce PLATO (Physical Language for Action through Tokenized Observations), a unified world-state representation for VLA perception and prediction. Learned from multi-view observations with geometric, semantic, and forward–inverse dynamics supervision, PLATO preserves 3D structure, semantics, and interaction-relevant information in a compact form. Continuous features or embeddings of discrete codes provide current-state context, and 128 discrete state tokens serve as autoregressive future-state targets alongside action supervision. Both roles leverage the VLM’s native conditioning and next-token prediction paradigm, without specialized dense prediction branches. Across VLA pretraining and post-training on three simulation benchmarks and four real-world tasks, PLATO improves manipulation success and robustness. Future-state supervision raises average real-world success by 11.7 and 25.4% under nominal and randomized conditions, respectively, and current-state context further improves viewpoint robustness, with a 40% gain over supervision alone in flower insertion. Beyond performance, our analyses show that PLATO reshapes VLA representations: geometry becomes more linearly accessible and RGB embeddings more consistent across viewpoints, revealing that explicit world-state inputs and targets substantially alter the physical structure learned by the VLA from robot data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.