acceptodds
Under review as a conference paper at ICLR 2027

Not All Evaluations Are Equal: Where Test-Time Compute Matters in World-Model Planning

Abstract

World-model planners spend their test-time compute scoring populations of imagined action sequences, and they spend it uniformly across the iterations of the search. We find that the planning decision is made early. In cross-entropy-method (CEM) planning with pretrained latent world models, late-round evaluations can be approximated or removed without a detectable change in success on PushT, Reacher, and Two-Room, whereas approximating early rounds can derail the search. This is consistent with the contraction of the proposal: as it narrows, a linear response operator fitted at run time around the proposal mean halves its held-out error on PushT and cuts it more than fourfold on Reacher. Substituting this operator into the early rounds lowers success significantly on PushT under CEM, iCEM, and MPPI and on Two-Room under MPPI, and raises the final cost on both tasks, whereas score noise of the same cost scale shows no such stage dependence on any of the three tasks. Reacher marks the boundary: its success is flat from 5 to 30 rounds, and early substitution does not hurt and even raises success. Whether early approximation hurts goes together with an offline quantity, the cost of stopping the native search early: the two agreed in all eight settings of a test with predictions fixed in advance. The finding can be used in two ways that trade time against fidelity: stopping two rounds early plans 12% faster at a small but significant increase in final cost, and Exact-to-Local CEM, which keeps early rounds exact and approximates late ones, has the same success outcome as native CEM on all 200 paired PushT decisions with 13% fewer exact queries, 3% less time, and no detectable increase in final cost. For the planners and world models we tested, late-round compute can be cut first; the early rounds are where approximation is risky.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.