Building Robust World Models for VLA RL: Visual Generalization and Long-Horizon Consistency
Abstract
Vision-language-action (VLA) policies can be post-trained in learned simulators, but current world models often fail under visual shifts and accumulate errors during long-horizon rollouts. We present Sword, a style-robust world model for policy post-training. Sword addresses these failure modes with two complementary components. Structure-Guided Style Augmentation (SGSA) uses depth, segmentation, and task-conditioned style transfer to diversify visual appearance while preserving task-relevant geometry and semantics. Dynamic Latent Bootstrapping (DLB) reuses cached model-predicted latents as context during fixed-window training, exposing the model to self-generated inputs without requiring full-sequence rollouts. We further introduce LIBERO-Mixed, an evaluation set combining original episodes with style-transferred episodes generated using prompts held out from training. Experiments on LIBERO show that Sword improves prediction quality and temporal consistency over representative action-conditioned world models, including under held-out style shifts. In GRPO post-training, using Sword as the simulator also yields higher VLA task success rates than using WoVR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.