From World States to Videos: V-JEPA for Driving Scene Generation
Abstract
Predicting a driving video is not only a problem of extending pixels: the generator must decide how traffic, road geometry, and camera motion evolve after the last observed frame. We ask whether a learned world state can make that evolution explicit. A compact Transformer predicts a future-oriented V-JEPA state from observed video, and a latent-video generator uses the state to synthesize the future. Neither future images nor motion metadata are available at inference. Across 130 unseen nuScenes scenes, predicted states improve a fixed semantic–geometric–motion measure by 6.88% over context-only generation in every seed, retaining more than 95% of the benefit of a teacher state extracted from the logged future. A matched predictive pathway built on VideoMAE uses the same predictor and generator budgets, yet predicted V-JEPA states perform 4.24% better. The gain spans all evaluator components and aggregation rules. Motion readouts explain why this is plausible: V-JEPA exposes temporal dynamics more clearly than VideoMAE on two scene cohorts. Oracle-conditioning experiments then show that the generator uses this information and degrades when state and scene no longer correspond. Together, these results support a context-to-state-to-video path in which predicted world states guide future-video synthesis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.