acceptodds
Under review as a conference paper at ICLR 2027

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Abstract

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the *apparent appearance* of generated videos to the *inherent reactivity* of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce **WorldExam**, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions with expected consequences left unstated and goal-directed behaviors with execution steps unspecified. Evaluation of 23 models shows that neither high visual quality nor strong control adherence guarantees inherent reactivity, while models with stronger reactivity can exhibit less reliable control. These findings suggest that progress requires combining these strengths, so that models both follow controls and generate plausible consequences implied by the scene.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.