acceptodds
Under review as a conference paper at ICLR 2027

Mental Simulation for Vision-Language Models on Intuitive Physics

Abstract

Vision-language models (VLMs) excel at interpreting what is visible, but can often struggle with future-dependent questions whose answers hinge on events that have not yet occurred. Such questions require anticipating how a scene will unfold—something humans do naturally through mental simulation. We propose a test-time approach that gives VLMs an analogous capability. To complement a VLM’s internal text-based reasoning, we use a generative video model to produce multiple plausible future rollouts conditioned on the observed video. These “mental simulations” are fed back to the VLM as additional visual context before it answers. We evaluate this simulate-then-reason framework on three intuitive physics benchmarks: Physion, CLEVRER, and real-world Physics-IQ videos. Across a diverse set of open-source and proprietary VLMs, augmenting models with imagined futures can improve prediction accuracy over both observed-only baselines and strong test-time language reasoning baselines, suggesting that allocating test-time computation to explicit video-based mental simulation can enhance multi-modal reasoning about the physical world. Video examples are available on our project page. (https://projectpagexxx.github.io/mental-simulation/)

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.