When Can a World Model Decide? Certifying Elite Decisions in Sampling-Based Planning
Abstract
Sampling-based planners such as CEM and MPPI turn a latent world model into a controller by trusting its ranking of candidate action sequences at every iteration. We show when that trust is warranted. A rollout error bound yields a decision certificate: the model places a candidate correctly in or out of the elite set whenever the candidate lies more than twice the error from the elite boundary. Executing every candidate a planner evaluates on Push-T, we find that the certified fraction falls from 0.67 to 0.29 over the search: the planner converges until its candidates differ by less than the model can resolve, and more search cannot recover the lost resolution. The certificate also names the remedy, a more accurate predictor or models with independent errors, and only the decisions at the elite boundary need it. We introduce EliteCert, in which the planning model screens every candidate and independently trained validator models re-score only a calibrated band around the K-th elite. EliteCert enters through the cost function, so CEM, iCEM and MPPI run unmodified. Under the official LeWM protocol it raises success by 9–14 points for all three planners on Push-T and by 2–3 points on Cube, beats the default configuration at 44% of its CPU time, matches re-scoring every candidate at a third of the extra compute, and adds a further 1.8 points to a planner whose predictor we retrained. An offline qualification test tells in advance which models can serve as validators. Finally, we prove and measure that gains on logged candidate pools do not transfer to the closed loop of an iterated planner, and evaluate every claim in the closed loop, with 1800 paired runs behind each headline comparison.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.