RoboAct-WorldVLA: Reinforcing Action Generation with a Latent Action-Understanding World Model
Abstract
Although Vision–Language–Action (VLA) models have shown strong potential in robotic manipulation, their imitation-driven nature still limits robust deployment under domain shifts. The core issue is not merely insufficient demonstrations, but the lack of inherent physical action understanding and the absence of a framework that uses such understanding to guide action generation. As a result, a VLA policy may produce actions that resemble demonstrations without knowing whether these actions move the system towards the goal or not. In this regard, we propose RoboAct-WorldVLA, a framework that reinforces action generation with a latent action-understanding world model. Given candidate actions from a flow-based generator or an off-the-shelf VLA policy, the world model predicts their future action representations in a pretrained disentangled latent space. A language-aligned action-value evaluator assesses how well the predicted future representation aligns with the prior action representations, suppressing uncertain or off-trajectory actions. The prior action representations are generated by an existing action representation encoder RoboAct-CLIP, which disentangles the inherent action dynamics from subject/object features for better generalizability. RoboAct-WorldVLA is first pretrained offline from demonstrations, failed trajectories, and counterfactual transitions. It is then fine-tuned online with environment rewards, which refine the action generator and calibrate the evaluator while preserving imitation, latent dynamics, semantic alignment, and uncertainty regularization. We evaluate RoboAct-WorldVLA on RLBench, ManiSkill3, and real-world manipulation tasks. Extensive experiments demonstrate that RoboAct-WorldVLA can consistently boost the performance of VLA backbones by 3.8–9.0 percentage points. Notably, the ablation study reveals the proposed method's active capability to determine required causal actions and evaluate current progress, guided by prior action representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.