Chain of Time: In-Context Physical Simulation with Image Generation Models
Abstract
We propose a novel cognitively-inspired prompting method to stabilize and interpret the physical simulation produced by autoregressive image generation models, called Chain-of-Time. Chain-of-Time involves generating a series of intermediate images during a simulation, and it is motivated by mental simulation in humans, as well as in-context reasoning in LLMs. Chain-of-Time is used at inference time, and requires no additional fine-tuning. We apply this method to synthetic and real-world domains, including 2-D graphics simulations and natural 3-D videos. These domains cover various physical dynamics, ranging from rigid-body rolling to harmonic oscillation and pulley systems. We found that Chain-of-Time simulation stabilizes the physics simulated by the state-of-the-art image generation models. Beyond examining performance, we also analyzed the specific states of the world simulated by an image model at each time step, which sheds light on the dynamics underlying these simulations. This analysis reveals insights that are hidden from traditional evaluations of physical reasoning, highlighting cases where image generation models can simulate physical properties such as velocity, gravity, oscillations, and pulley systems. Like humans, image generation models are able to accurately predict future world states over short time periods. By providing a mechanism analogous to working memory, Chain-of-Time enables a model to stitch together a series of short-term predictions into an accurate long-term simulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.