Seeing Ahead: Plug-in World Models for Predictive Reasoning in Frozen VLMs
Abstract
Understanding a video involves recognizing the visible scene and anticipating how it may change. Video world models learn predictive representations; vision–language models (VLMs) connect visual evidence to questions and answers. We study how to combine these capabilities while retaining the VLM's own perception. We propose FoReST (Forecast Residual State Transfer), which adds predictive updates to a frozen VLM's visual tokens. One variant translates changes predicted by a pretrained world model; the other forecasts the update from that model's observed-video features. Only a compact head is trained from video continuations, without question–answer supervision. We introduce CLEVRER-Next, a next-event QA dataset built from existing CLEVRER videos and annotations, with explicit observation cutoffs and retained continuations for diagnosis. On CLEVRER-Next and ComPhy, latent transfer improves accuracy by 11.5 and 11.7 percentage points over the same frozen VLM given observed video. Against a trained forecaster of the VLM's own features, latent transfer provides additional gains on ComPhy, while direct forecasting improves CLEVRER-Next collision prediction. Controlled comparisons show that retaining native vision reduces degradation from imperfect forecasts and that prediction-input training can improve their translation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.