Scaling Vision-Language-Action Models with Video Generative Priors
Abstract
Scaling vision-language-action (VLA) models requires substantial robot data and computation. Pretrained video generators provide visual and temporal representations, but rendering future frames adds iterative sampling and re-encoding. We introduce VID-Policy, a video-generation-model-enhanced policy that transfers history-conditioned representations from a frozen video generator into a pretrained VLA. A projector maps these features to policy-side keys and values; gated additive attention combines them with native context while preserving the action generator. We assess source precision, history content, and closed-loop control in two policy families. In a fixed-seed, 24-task, five-episode-per-task DIAL development screen, BF16 video features yield 80.8% success, 4.2 percentage points above our reproduced Native DIAL reference. On the full 24-task, 50-episode-per-task RoboCasa-GR1-Tabletop benchmark, WALA-based VID-Policy reaches 72.33% success, improving over our WALA re-implementation by about 1.83 percentage points. The results show that video representations can guide action without sampling future frames.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.