WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
Abstract
Reinforcement learning (RL) can improve Vision-Language-Action (VLA) policies beyond supervised fine-tuning, but critic-based methods depend on value estimates that must account for partial observability. Critics operating on current-observation features can miss temporal information relevant to task progress, while simply adding observation history does not reliably improve policy performance in our experiments. We propose the World Critic Model (WCM), a history-conditioned critic that combines value estimation with action-conditioned latent prediction. Built on a lightweight LeJEPA-based architecture, WCM encodes a short observation history and a language instruction into a shared temporal representation, from which it estimates returns and predicts the next observation latent under the executed action. The prediction objective complements return supervision by encouraging the representation to retain transition-relevant information, without requiring pixel reconstruction or imagined rollouts for value estimation. WCM integrates into both on-policy and off-policy RL pipelines and supports diverse VLA backbones. Experiments spanning 149 tasks across four simulation benchmarks demonstrate improvements in policy performance and generalization, with ablations supporting the benefit of combining history with latent prediction. Evaluations on seven real-world manipulation tasks further demonstrate the effectiveness of WCM in off-policy VLA reinforcement learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.