Reasoning-Structured Videos: A Stratified Diagnostic Suite for Compositional Consistency in World Models
Abstract
Existing evaluations of video world models measure visual fidelity, control adherence, and revisit consistency. We extend this perspective from within-rollout revisitation to agreement across distinct histories reaching the same reference state. We introduce Reasoning-Structured Videos, an Unreal Engine 5 benchmark for evaluating these relationships in static scenes. The dataset comprises 11,377 exploration and relation-structured videos with action and camera-pose annotations. Its relation-structured trajectories form rooted graphs with shared initial states and annotated endpoint relationships. Inverse tests recovery after retracing a path, Loop tests return after an exploratory detour, and Equivalence tests agreement between distinct paths reaching the same state. Our evaluation separates self-consistency from reference fidelity and complements image-based metrics with layout and geometry readouts. Across nine world-model configurations, return-consistency rankings broadly align with distributional quality, whereas cross-path agreement yields different rankings. Strong layout agreement can coexist with poor geometric fidelity, and revisit consistency deteriorates over longer intervals. Together, these findings show why cross-history agreement, reference fidelity, and long-range retention should be evaluated separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.