acceptodds
Under review as a conference paper at ICLR 2027

DARC: Decision-Aware Readout Calibration for Generative World Models

Abstract

Generative world models are increasingly used to evaluate candidate futures for planning and control, yet better numerical inference does not necessarily produce better downstream decisions. We study this mismatch by separating three notions of fidelity: numerical fidelity of the sampler, environmental fidelity to the underlying dynamics, and decision fidelity after a planner-specific readout. Across 13 DeepMind Control environments, increasing sampler refinement from 1 to 32 steps improves sample sharpness while worsening simulator-referenced PSNR and increasing planner readout error, showing that numerical refinement can amplify errors that matter to the decision. We introduce DARC, a decision-aware calibration framework for frozen generative world models. DARC uses simulator-referenced attribution to separate finite-inference error from learned-model error, then applies a readout-specific, mean-preserving dispersion transport that calibrates the statistics actually consumed by the downstream planner without retraining the world model. Across 135 latent world models and 1,620 evaluation cells, DARC removes 25.0% of planner readout error on average, with a 95% confidence interval of [16.3%, 34.5%], while harming only 1/1,620 cells. In contrast, latent-variance calibration yields −16.5% error removal and harms 330/1,620 cells. These results show that world-model inference should be evaluated and calibrated through the downstream decision functional, rather than through generative or numerical fidelity alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.