MetaReach: Meta-Reachability Group Relative Policy Optimization for Video World Models
Abstract
Conditional video world models predict future visual states under given actions, but post-training video world models with group relative policy optimization (GRPO) remains challenging. Existing methods freeze the decoder and compute rewards against ground-truth frames, introducing an irreducible residual that distorts candidate ranking because the decoder's reachable space is limited (target-set mismatch). They also broadcast a single sequence-level advantage to all token blocks, ignoring the causal structure of autoregressive generation (long-horizon credit misallocation). We propose MetaReach, a bi-level meta-learning framework that jointly addresses both issues. The inner loop adapts the decoder to reconstruct frames within the world model's output distribution, and the outer loop updates the world model under the adapted decoder. This reachability-calibrated meta-learning (RCM) eliminates target-set mismatch at its source. We further introduce autoregressive credit assignment (ACA), which assigns a position-specific advantage to each token block according to its future horizon. Gradient projection preserves the original objective. Experiments on RT-1, RoboNet, and VP2-RoboSuite on a 15-frame prediction horizon show that MetaReach consistently outperforms GRPO-based baselines and recent long-horizon post-training methods, reducing LPIPS by 7.7% and MSE by 13.3% over RLVR-World on RT-1.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.