Visual Fidelity and Physical Reliability in World-Model Rollouts
Abstract
World models are often judged by predicted observations, even when autonomous rollouts are used to reason about physical state. We ask how much a model's native visual prediction error tells us about physical accuracy at the rollout level. Across DreamerV3 and PlaNet, lower visual loss selects rollouts with lower autonomous physical readout error than matched random selection over a broad coverage range, but this ranking does not establish a numerical physical-error scale. We therefore ask what visual loss misses. Around some rollout roots, the queried physical quantity changes strongly under perturbations that produce comparatively little change in the rendered trajectory. A finite-scale query-to-render sensitivity score predicts realized perturbation responses and autonomous physical error beyond visual loss in two independent DreamerV3 cohorts. In a prospectively specified third cohort of 500 episode families, using the same score after visual filtering reduces mean physical error from 0.198 to 0.166 at matched 10% coverage, a 16.0% relative reduction (95% whole-family bootstrap CI for the mean difference, [-0.0488, -0.0162]). Finally, we show that positive-loss screening results do not determine whether physical error must vanish as visual loss approaches zero. Visual fidelity is useful for ranking world-model rollouts, but is not sufficient evidence that a particular physical query is reliable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.