How Many Samples Should a Planner Draw? The Answer Depends on a Convention Nobody Reports
Abstract
Sampling-based planners (CEM, MPPI) are the deployment-time engine of modern world-model agents, and a practitioner's first question about them is how to spend planner compute: draw more candidate action sequences, or refine them over more iterations? We show that the first half of this question is not well posed as usually asked. Implementations differ in whether the elite set is a fixed fraction of the population (mbrl-lib's default, used by PETS and PlaNet) or a fixed count (TD-MPC2's convention), and this choice, usually inherited from a library default, can reverse the sign of the measured effect of drawing more samples. On identical PlaNet models, seeds and episodes, increasing the population changes return by under one convention and under the other (both ); the other two PlaNet tasks point the same way, while on PETS the convention changes the effect's size (significantly only on Ant) but not its sign. Instrumenting the optimizer gives a post-hoc mechanism: the elite count sets how concentrated the refit proposal is, so tying it to the population widens the proposal as sampling grows (PETS dispersion ), while a fixed count tightens it (). Sample-count comparisons are therefore not commensurable unless the elite convention is stated. The second half of the question holds the population fixed, so the convention does not enter: across released TD-MPC2 agents on 14 tasks in two domains and two further agent families, refinement iterations are the more reliable way to convert planner compute into return. They recover full-planner performance on weak-prior high-dimensional bodies and on pixel PlaNet, where width does not (e.g. , , where the best width-only configuration still falls short), with 90% of the gain arriving between the second and fourth iteration on all four TD-MPC2 tasks swept (4-D to 38-D); on state-based PETS one well-configured wide pass matches them, and on a poorly fit model they hurt. How much refinement a task needs tracks the prior-to-task gap, a contrast carried by three weak-prior high-dimensional bodies. Every sweep was preregistered with dated amendments; of the twenty predictions tested, four held, twelve failed, one held in part and three could not be established, and we report all of them.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.