acceptodds
Under review as a conference paper at ICLR 2027

SpaceTimeBench: Evaluating Space And Time In The World Model

Abstract

Generative world models aim to serve as universal physical simulators, yet existing benchmarks fail to verify the fundamental prerequisite of Newtonian Law: that space and time must remain absolute and mathematically consistent across different observer perspectives. To bridge this evaluation gap, we introduce SpaceTimeBench, a novel Video-to-Video framework that utilizes prompt-injected geometric anchors alongside DePT3R and SAM 2 to extract strictly normalized 4D coordinates, testing whether models perfectly conserve spatial distances and temporal durations across varied camera views. Extensive evaluations across 17 leading architectures reveal a severe capability deficit quantified by our Final Spacetime Yield (FSY), demonstrating that even state-of-the-art models peak at a mere 65.0% success rate. Crucially, SpaceTimeBench functions as an evaluation oracle to expose a False Discovery Rate (FDR) of up to 42.9% regarding spurious compliance in existing benchmarks, proving that evaluating Newton's First Law without verifying underlying spatiotemporal prerequisites severely inflates scores and masks the reality that current architectures remain mere two-dimensional pattern matchers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.