acceptodds
Under review as a conference paper at ICLR 2027

Building Robust World Models for VLA RL: Visual Generalization and Long-Horizon Consistency

Abstract

Vision-language-action (VLA) policies can be post-trained in learned simulators, but current world models often fail under visual shifts and accumulate errors during long-horizon rollouts. We present Sword, a style-robust world model for policy post-training. Sword addresses these failure modes with two complementary components. Structure-Guided Style Augmentation (SGSA) uses depth, segmentation, and task-conditioned style transfer to diversify visual appearance while preserving task-relevant geometry and semantics. Dynamic Latent Bootstrapping (DLB) reuses cached model-predicted latents as context during fixed-window training, exposing the model to self-generated inputs without requiring full-sequence rollouts. We further introduce LIBERO-Mixed, an evaluation set combining original episodes with style-transferred episodes generated using prompts held out from training. Experiments on LIBERO show that Sword improves prediction quality and temporal consistency over representative action-conditioned world models, including under held-out style shifts. In GRPO post-training, using Sword as the simulator also yields higher VLA task success rates than using WoVR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.