acceptodds
Under review as a conference paper at ICLR 2027

Planning Is Test-Time Scaling: A World Planning Model

Abstract

Effective decision-making requires more than predicting the future. An agent must also know how to use predicted outcomes to improve its actions. We introduce the World Planning Model (WPM), a learned iterative decision process that couples future prediction with action refinement in a single network. Starting from an initial action proposal, WPM predicts its consequences and uses the resulting future representation to update the action, repeatedly applying the same operator to produce progressively refined decisions. By training prediction and refinement jointly, WPM learns not only task-relevant dynamics, but also how predictive information should be translated into action updates. This replaces explicit candidate sampling, scoring, and selection with implicit refinement, while retaining the ability to allocate additional computation at test time through more refinement steps. In offline goal-conditioned control from pixels, WPM matches or exceeds explicit planners built on the same model on most environments. Furthermore, WPM's success rises with refinement depth on most environments without retraining. Controlled ablations further indicate that predictive supervision improves the learned decision process. These results suggest a general alternative to both one-shot policy execution and explicit search: learning a reusable planning operator whose computation can be scaled at test time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.