EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
Abstract
Chunked vision–language–action (VLA) policies generate multi-step actions from one observation and typically re-observe after executing the action chunk. Earlier actions change the scene conditions for later steps, while occlusion can compromise visual feedback. The previous scene estimate is typically not propagated to the next control call. Although spatial and temporal VLAs enhance geometric modeling and maintain temporal memory, respectively, their internal representations may become stale as robot actions alter object poses and contact states. We introduce EvoScene-VLA, which uses compact scene tokens to unify within-chunk scene prediction, cross-chunk state propagation, and observation-based correction. The action expert jointly denoises actions and future scene states in a single flow-matching process, allowing predicted scene changes to inform action generation. The final scene state serves as the next prior for subsequent correction. At the next control call, the vision–language model (VLM) fuses the incoming visual input with this prior. To preserve control-relevant geometry, two-level geometric anchoring applies local depth supervision to per-view scene features and global 3D feature supervision to the fused scene state; a shared geometric decoder imposes the same constraints on current and future states. Since demonstrations do not directly provide future scene states, we train a scene predictor using 3D teacher features from future images. Its predictions provide targets for the action expert's future scene states. At deployment, we remove both Scene Predictor and Geometric Anchor. EvoScene-VLA improves average success over the matched baseline by 2.4 and 2.2 percentage points on 31 RoboTwin tasks under fixed and randomized settings, respectively, and by 2.7 points on LIBERO. Across three real-robot tasks, average success rises from 37.3% to 42.0%. Ablations show that a recurrent scene state alone does not necessarily improve performance, whereas state propagation can improve control when combined with geometric and future scene supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.