acceptodds
Under review as a conference paper at ICLR 2027

ChronoSpatial: A Procedural Video Benchmark for Memory and World-State Reasoning

Abstract

Understanding a video requires maintaining a changing world state: an object may move, disappear behind an obstacle, change hands, or be revisited from a different viewpoint. We introduce ChronoSpatial, a procedural benchmark for evaluating these abilities through delayed, four-choice questions about rendered event sequences. The initial collection contains 215 videos spanning sixteen scenario families, including household activities, navigation, object transfers, and causal interactions. Executable scene descriptions provide reproducible generation and simulator-derived answers. An isolated evaluation harness exposes only anonymous videos and questions, supporting both direct vision-language inference and an agent that inspects video through a bounded visual tool. On the 195-item evaluation split, the GPT-6 Astra agent achieves 79.0%, 96.9%, and 97.4% accuracy with budgets of 16, 64, and 256 inspected frames. Two completed open-model baselines, MiniCPM-V-4.6 and GLM-4.6V-Flash, achieve 28.7% and 25.6% at 64 sampled frames under a separate direct-inference protocol. These initial results motivate closer examination of evidence access, response handling, and sequential complexity. ChronoSpatial provides a reproducible platform for studying visual state tracking while making differences in inference interfaces and computational budgets explicit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.