acceptodds
Under review as a conference paper at ICLR 2027

StopTheDream: Outcome-Grounded Candidate Comparison for Adaptive Search in World-Model Planning

Abstract

A world-model planner can always select the best candidate it has imagined, even when its predictions provide weak evidence that the selected action will achieve the goal. Generating more candidates can help, but adds computation and can amplify the same scoring error. We introduce StopTheDream, a framework that connects candidate actions to their predicted consequences and uses this evidence for candidate comparison and bounded search. The procedure makes generation, comparison, and computation allocation explicit, grounding its scoring objectives in observed actions and outcomes. On 120 hierarchical world model (HWM) maze tasks, the policy solves 117 tasks versus CoVO-MPC’s 113, using 48.58% fewer latent transitions and 24.38% less planning time. The original HWM planner solves 105 tasks with the same world model. On 1,000 CAST visual trajectories, it attains 1.3629 absolute trajectory error, 4.18% below a locally trained PiJEPA baseline, with 54.11% fewer predictor calls. These results show higher observed planning quality with substantially less predictive computation, supporting joint evaluation of decision quality and computational cost

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.