StopTheDream: Outcome-Grounded Candidate Comparison for Adaptive Search in World-Model Planning
Abstract
A world-model planner can always select the best candidate it has imagined, even when its predictions provide weak evidence that the selected action will achieve the goal. Generating more candidates can help, but adds computation and can amplify the same scoring error. We introduce StopTheDream, a framework that connects candidate actions to their predicted consequences and uses this evidence for candidate comparison and bounded search. The procedure makes generation, comparison, and computation allocation explicit, grounding its scoring objectives in observed actions and outcomes. On 120 hierarchical world model (HWM) maze tasks, the policy solves 117 tasks versus CoVO-MPC’s 113, using 48.58% fewer latent transitions and 24.38% less planning time. The original HWM planner solves 105 tasks with the same world model. On 1,000 CAST visual trajectories, it attains 1.3629 absolute trajectory error, 4.18% below a locally trained PiJEPA baseline, with 54.11% fewer predictor calls. These results show higher observed planning quality with substantially less predictive computation, supporting joint evaluation of decision quality and computational cost
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.