acceptodds
Under review as a conference paper at ICLR 2027

Beyond Visual Fidelity: Comparing Categorical and Continuous World Models

Abstract

Visual world models can now generate compelling interactive worlds, but visual realism does not guarantee faithful dynamics. To understand how modeling choices affect world model behavior, we present a systematic comparison that disentangles prediction representation from generation procedure, choices often bundled in approaches such as autoregressive token models and whole-frame flow matching. In controlled 2D and 3D environments, we evaluate dynamics validity, calibration, and visual quality. Failures such as disappearing game characters echo those observed in more complex environments, but here their impact can be measured against known dynamics. With lossless visual encoding in 2D, discrete- token models achieve higher rollout validity at lower training compute, particularly for stochastic entities. However, high validity can coexist with poor calibration: generating permissible trajectories does not imply reproducing their probabilities. In 3D, visual degradation from lossy tokenization does not consistently translate into worse long-horizon metrics. These results show how representation and decoding choices affect validity and calibration differently, demonstrating the importance of evaluating world models beyond visual fidelity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.