WorldRoamBench: Evaluating Long-Horizon Stability of Interactive World Models
Abstract
Despite rapid progress in interactive world models (IWMs), their long-horizon stability under sustained action inputs remains insufficiently evaluated. We introduce WorldRoamBench, an open-world benchmark addressing this gap across four complementary dimensions: (i) Action: pose-based discrete-action accuracy reducing sensitivity to cross-model semantic scale disparity and exposing frame-level errors hidden by trajectory scores; (ii) Vision: sliding-window drift metric capturing non-monotonic mid-sequence collapse missed by start-vs-end comparisons; (iii) Physics: evaluation of physical plausibility across mechanics, optics, and 3D consistency, gated by camera-motion and subject-tracking checks; (iv) Memory: a trajectory-aware protocol reducing confounding from action-following errors, evaluating scene memory via transition-localized 3D point-cloud reconstruction and subject memory via tracking-plus-VLM reasoning. The benchmark comprises 1000+ test cases across Nature, Urban, and Indoor scenes in first/third-person views with WASD 10-60 s continuous interaction. Evaluating 12 open/closed-source models reveals none reliably satisfies all dimensions; even the best long-horizon first-person overall score is only 67.17/100 (Genie 3). Advances on WorldRoamBench are steps toward IWMs that are stable, physically grounded, memory-faithful, and deployable in real-world applications.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.