ChronoSpatial: A Procedural Video Benchmark for Memory and World-State Reasoning
Abstract
Understanding a video requires maintaining a changing world state: an object may move, disappear behind an obstacle, change hands, or be revisited from a different viewpoint. We introduce ChronoSpatial, a procedural benchmark for evaluating these abilities through delayed, four-choice questions about rendered event sequences. The initial collection contains 215 videos spanning sixteen scenario families, including household activities, navigation, object transfers, and causal interactions. Executable scene descriptions provide reproducible generation and simulator-derived answers. An isolated evaluation harness exposes only anonymous videos and questions, supporting both direct vision-language inference and an agent that inspects video through a bounded visual tool. On the 195-item evaluation split, the GPT-6 Astra agent achieves 79.0%, 96.9%, and 97.4% accuracy with budgets of 16, 64, and 256 inspected frames. Two completed open-model baselines, MiniCPM-V-4.6 and GLM-4.6V-Flash, achieve 28.7% and 25.6% at 64 sampled frames under a separate direct-inference protocol. These initial results motivate closer examination of evidence access, response handling, and sequential complexity. ChronoSpatial provides a reproducible platform for studying visual state tracking while making differences in inference interfaces and computational budgets explicit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.