Mental Simulation for Vision-Language Models on Intuitive Physics
Abstract
Vision-language models (VLMs) excel at interpreting what is visible, but can often struggle with future-dependent questions whose answers hinge on events that have not yet occurred. Such questions require anticipating how a scene will unfold—something humans do naturally through mental simulation. We propose a test-time approach that gives VLMs an analogous capability. To complement a VLM’s internal text-based reasoning, we use a generative video model to produce multiple plausible future rollouts conditioned on the observed video. These “mental simulations” are fed back to the VLM as additional visual context before it answers. We evaluate this simulate-then-reason framework on three intuitive physics benchmarks: Physion, CLEVRER, and real-world Physics-IQ videos. Across a diverse set of open-source and proprietary VLMs, augmenting models with imagined futures can improve prediction accuracy over both observed-only baselines and strong test-time language reasoning baselines, suggesting that allocating test-time computation to explicit video-based mental simulation can enhance multi-modal reasoning about the physical world. Video examples are available on our project page. (https://projectpagexxx.github.io/mental-simulation/)
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.