acceptodds
Under review as a conference paper at ICLR 2027

LongWM-Bench: Benchmarking Persistent and Navigable Worlds over Long-Horizon Interaction

Abstract

Short-term visual quality and action following offer limited evidence of whether an interactive world model remains reliable during extended use. We introduce LongWM-Bench, a benchmark of 300 tasks with target rollout durations of 60–180 seconds. Two task suites evaluate sustained control and revisit memory through combined inputs, repeated changes in control, and six families of trajectories with delayed and repeated revisits. We combine automatic visual and motion measurements with checklists tailored to each scene to assess visual quality, temporal consistency, revisit retention, physical and causal validity, and camera responses to actions. Evaluating 12 interactive world models reveals distinct capability profiles: models with similar appearance consistency differ substantially in revisit retention, while accurate camera responses can coexist with weak physical and causal validity. Coherent scene evolution remains a common weakness under our evaluation. These findings highlight the need to evaluate whether generated worlds remain controllable, retain previously observed content, and evolve coherently throughout interaction. Evaluation code and example data are included in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.