acceptodds
Under review as a conference paper at ICLR 2027

From Signal Fidelity to Future-State Alignment: A Hierarchical Evaluation of Video World Models

Abstract

Video generation models are increasingly viewed as candidate world models, yet visually convincing rollouts do not necessarily predict how a depicted scene actually evolves. Existing evaluations often conflate frame-level signal similarity or no-reference video quality with predictive state understanding. This work introduces FPAbench (Fidelity, Plausibility, and Alignment benchmark), a hierarchical evaluation of video world models that separates three questions: Signal Fidelity (L1), measuring frame and motion-signal proximity to an observed future; Video Plausibility (L2), measuring the visual coherence, temporal smoothness, motion, and apparent physical validity of an unconstrained rollout; and Future State Alignment (L3), measuring whether the same rollout preserves entities, trajectories, spatial interactions, and causal state evolution relative to the observed future. FPAbench evaluates 15 video generation systems on 2,000 five-frame sequences spanning 13 physical-event categories. Across every evaluated system, L3 is lower than L2, and the same pattern holds in 190 of 195 model-category comparisons. The gap is largest for state-transition events involving gravity, collisions, fluids, deformation, and precarious configurations, whereas relatively stable scenes are easier. Motion dynamism is not predictive of future-state alignment. These results show that signal fidelity and plausible video appearance are necessary but insufficient evidence of world-model capability, motivating hierarchical evaluation from signals to event-specific future states in observed scenes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.