acceptodds
Under review as a conference paper at ICLR 2027

WorldProbe: When Video World Models Fall Apart under Long-Horizon Spatial Exploration

Abstract

Video world models generate future visual observations from camera or action inputs, with the goal of synthesizing persistent 3D worlds that can be explored across time and space. Existing benchmarks typically evaluate short interactions or rely on camera trajectories fixed before generation. The performance of video world models under long-horizon exploration with extensive camera motion remains under-explored. We introduce WorldProbe, a diagnostic benchmark that uses a goal-conditioned multimodal agent to adapt camera actions online during generation. WorldProbe comprises 164 test cases across four probing families. Across 10 state-of-the-art models, we generate more than 1,600 rollouts totaling approximately 800 minutes, with horizons up to 180 seconds. Our evaluation across four dimensions reveals both the degradation of local measures and the emergence of long-range inconsistencies. We find that recent models sustain visual and content quality for longer, while long-range memory and evolution remain weak. These results expose a gap between locally stable generation and persistent world consistency that is largely missed by existing benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.