Trust Before Use: When Do World-Model Quality Metrics Predict Control?
Abstract
World models are being scaled to billions of parameters and fine-tuned on robot data in the hope of unlocking new skills, yet it remains unclear which properties of a world model make it useful for downstream control. Practitioners typically judge world models by prediction loss or video fidelity, and learn whether a model is useful only after training and evaluating a policy, a slow and costly loop. We ask when world-model quality metrics predict control. We find that prediction loss and video fidelity can mislead: a world model with lower prediction loss can yield a worse policy, and a world-action model with far worse video fidelity can control without a detectable loss in success. We study the expected signal-to-noise ratio (ESNR) of policy gradients computed through a world model, which measures the gradient signal the model offers to policy extraction and can be computed during training without a trained policy or environment interaction. Across policy-extraction methods and world-model architectures on DMControl and Meta-World, ESNR tracks downstream performance more closely than prediction loss, and it tracks the planning performance of large pretrained world models at a fraction of the evaluation cost. On a 6B-parameter world-action model trained by behavior cloning, a gradient signal-to-noise ratio of the action loss orders the four model variants exactly as closed-loop success does, whereas LPIPS and SSIM rank the most successful variant last. Visualizations and code will be available at https://trust-before-use.github.io.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.