acceptodds
Under review as a conference paper at ICLR 2027

Sanity Checks for Planning with Video World Models

Abstract

Video world models support planning by predicting the outcomes of candidate plans. We examine how strongly those predictions depend on the plan and how consistently the evaluator selects it. We introduce three checks that compare plan changes with sampler noise, compare predictions with simulator futures under alternative plans, and repeat the evaluator’s decision on identical inputs. In World-in-World image-goal navigation with Wan2.2, we analyze 33 episodes and 536 replanning steps. Changing one action affects the goal-similarity score less than changing the sampler seed: the ratio is 0.81 (95% CI [0.69, 0.95]) for the VLM planner and 0.57 ([0.54, 0.60]) for the heuristic planner. Predictions favor their own plan’s true future over perturbed plans’ futures by a margin equal to 1.5% of prediction error. At 43% of evaluated steps, none of five repeated VLM queries reproduces the executed plan. Four training-free confidence signals have episode-success AUCs of 0.47–0.66. Prediction fidelity has the opposite association (AUC = 0.12), with more accurate predictions in failed episodes. These measurements distinguish action sensitivity, decision consistency, and task success, and provide a practical protocol for evaluating confidence signals in video-based planners.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.