How Far Should an Agent Imagine? Decision-Stable Planning with Learned World Models
Abstract
Model-based reinforcement learning uses learned world models to evaluate future consequences, yet deeper imagination brings both delayed decision information and accumulating model error. Existing methods often adapt rollout depth using prediction uncertainty alone, although uncertainty matters only when it can change the preferred action. We introduce MarginPlan, which selects planning depth according to decision stability. For each action, it constructs uncertainty-aware return intervals over plausible dynamics models and identifies whether the same action remains optimal across them. Planning is declared decision-stable when the lower return bound of the preferred action exceeds the upper bounds of all alternatives. MarginPlan extends imagination only when the expected reduction in action ambiguity justifies the additional model-induced uncertainty. We evaluate controlled tasks that independently vary model uncertainty, action-value margins, and delayed rewards, including matched-uncertainty settings requiring different planning depths, together with computation-matched baselines. This formulation shifts adaptive planning from uncertainty-aware rollout truncation to decision-stable imagination, where planning depth is determined by whether additional lookahead can reliably improve the action decision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.