CalibWM: Action-Fidelity Calibrated Uncertainty for Frozen Video World Models
Abstract
Action-conditioned video world models support robot policy evaluation and imagination-based planning, but their training and evaluation focus largely on vi- sual fidelity. Recent benchmarks show that photorealistic rollouts can misrepre- sent the effects of executed actions and depict failures as successes. Uncertainty estimates calibrated to visual correctness inherit this mismatch. We introduce CalibWM, a lightweight probe on a frozen action-conditioned video world model that produces a spatiotemporal uncertainty map aligned with action-fidelity fail- ure, defined by whether a frozen inverse dynamics model can recover the executed action from generated frames. A quantile objective regularizes uncertainty to grow with rollout horizon, while a spatial term localizes the map and an alignment term trains its readout to predict action-fidelity failure. Per-horizon, segment- level split conformal calibration provides finite-sample coverage for aggregated normalized prediction-error scores under exchangeability, with unit-level impli- cations determined by the spatial aggregator. On Bridge and DROID, CalibWM achieves AUROCs of 0.782 and 0.761, respectively, compared with 0.694 and 0.679 for the strongest baseline trained without the action-fidelity label. Its mar- gins over the strongest control trained with that label are +0.020 (clustered 95% CI [0.009, 0.031]) on Bridge and +0.017 ([0.006, 0.028]) on DROID. On a human- annotated optimism-bias subset with uniformly high visual quality and rare fail- ures, CalibWM achieves 0.634 AUPRC against a chance level of 0.09; an oracle visual proxy with access to the ground-truth future reaches 0.106. We establish the coverage result at each fixed horizon. Under ordering consistency, we also prove that the action-fidelity readout has AUROC no lower than a visual proxy computed from the generated rollout alone, provided that the proxy carries no ad- ditional information about action-fidelity failure given the probe’s feature context.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.