Latent Lookahead for Action Conditioned Predictive Structure in Frozen Vision Language Models
Abstract
Pretrained vision–language models (VLMs) are increasingly reused as internal representations for multimodal agents, yet it remains unclear whether states optimized for understanding preserve information about how visual evidence changes after an action. We ask whether question- and history-conditioned representations from a frozen VLM contain recoverable, local action-conditioned predictive structure. Latent Lookahead uses a shallow readout and a low-capacity transition model to predict future latent states, then ranks candidate actions through short latent rollouts while leaving the backbone fixed. Across three VLMs and image, document, and video decision environments, learned transitions outperform persistence and action-independent predictors. The advantage persists over short open-loop rollouts, but weakens with horizon, policy-induced distribution shift, and poorer training coverage. From identical decision states, predicted futures increase best-realized-action match from 55.9% to 64.3% and reduce action regret from 0.0283 to 0.0146 relative to direct scoring. Latent Lookahead also improves task performance by about two points on average with roughly 30% additional compute and no extra backbone calls per real interaction step. These results show that local predictive information can be recovered from frozen VLM representations and used for short-horizon visual evidence selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.