Lazy Belief Expansion: Deciding What to Observe for World-Model Planning
Abstract
World models predict action consequences, but prediction alone does not determine which evidence the next decision needs. Exposing all available information adds processing cost and can distract reasoning, while a fixed compressed view can omit a decision-critical detail. We introduce *Lazy Belief Expansion* (*LBE*), a mechanism that lets the agent decide which evidence should condition world-model predictions and the next action. *LBE* retains source access while maintaining a compact working belief that can expand as evidence needs emerge. A controller chooses whether to *READ* from the current source, *ACT* to obtain missing evidence through interaction, or *COMMIT* the belief to task-directed planning; both environmental routes use the same planner. In a training-free web instantiation using *WMA* as the base planner, calibrated success rates for the 812-task benchmark are 19.33% for *LBE* and 18.72% for *WMA*. Among successful *LBE* tasks, final planning evidence is reduced by 54.6% relative to the corresponding untruncated source observations and by 52.8% relative to the context-capped full-observation inputs supplied to *WMA*; for tasks successfully completed in at most 5, 6–10, and 11–15 steps, *LBE*'s mean input-plus-output token usage is 22.4%, 16.0%, and 4.7% lower than *WMA*'s, respectively. Together, these results show that *LBE* integrates deciding what to know into deciding what to do without modifying the underlying predictive world model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.