acceptodds
Under review as a conference paper at ICLR 2027

Mind the Space: Benchmarking Spatial State Evolution in Video World Models

Abstract

A generated video can depict the requested action while failing to realize its spatial consequences. We introduce MTS-Bench (Mind the Space Benchmark), which evaluates this distinction through interventions whose task-supported spatial consequences are not fully narrated in the prompt. Its 337 curated cases pair image-to-video (I2V) and text-to-video (T2V) formulations across scene, view, and coupled motion. Case-specific rubrics assess initial-state fidelity, instruction execution, consequence prediction, temporal consistency, and endpoint completion. We evaluate nine I2V and eight T2V generators on a fixed 150-case subset, using a VLM scorer audited against blinded human ratings and selectively calibrated by dimension. The highest-scoring generator obtains 80.2 on instruction following and 56.2 on consequence prediction in I2V; the corresponding T2V scores are 77.2 and 62.4 (out of 100). Matched qualitative examples expose action without the required spatial change, while dimension profiles distinguish generators with similar aggregate scores. Evaluator analysis further shows that strong agreement with aggregate rankings can coexist with weak discrimination between videos. MTS-Bench makes intervention-induced spatial evolution an explicit, diagnosable evaluation target for video world models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.