acceptodds
Under review as a conference paper at ICLR 2027

Roz: A Taste Test for World Model Latents

Abstract

State of the art latent dynamics world models achieve impressive performance for robot planning. But do they encode enough information to accurately unfold dynamics? In this paper, we show that popular world models today achieve strong results on tasks that reduce to image similarity but forget key physical details like mass, deformation, friction, and bounciness. We achieve this by defining and justifying commonsense properties for what ought to be encoded in such a world model and introducing two new scalable evaluation methods. Our evaluation, which we call Roz, totals 1734 individual scenarios. We run Roz on a series of models and validate our eval on control tasks. On our own model, AnonWM, we find that performance roughly improves with training steps, while not necessarily being correlated with reconstruction MSE. We argue current model failures are a symptom of evaluation strategies (probes and planners) that are opaque, model-specific, and high-compute. Roz requires only a model that can ingest still frames or video, a way to extract the latents it predicts from, and a distance function between them, and it can be run in one pass. We hope future evaluations will adopt Roz's evaluation techniques. More broadly, we hope Roz can be a case study for a future where benchmarks allow rapid iteration and comparison across latent world models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.