LongWM-Bench: Benchmarking Persistent and Navigable Worlds over Long-Horizon Interaction
Abstract
Short-term visual quality and action following offer limited evidence of whether an interactive world model remains reliable during extended use. We introduce LongWM-Bench, a benchmark of 300 tasks with target rollout durations of 60–180 seconds. Two task suites evaluate sustained control and revisit memory through combined inputs, repeated changes in control, and six families of trajectories with delayed and repeated revisits. We combine automatic visual and motion measurements with checklists tailored to each scene to assess visual quality, temporal consistency, revisit retention, physical and causal validity, and camera responses to actions. Evaluating 12 interactive world models reveals distinct capability profiles: models with similar appearance consistency differ substantially in revisit retention, while accurate camera responses can coexist with weak physical and causal validity. Coherent scene evolution remains a common weakness under our evaluation. These findings highlight the need to evaluate whether generated worlds remain controllable, retain previously observed content, and evolve coherently throughout interaction. Evaluation code and example data are included in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.