IMWM: Intuition Models Complement World Models for Latent Planning
Abstract
Planning with a learned latent world model is a promising route to control from raw pixels, and when such a planner fails, the usual remedy is a better world model. We show experimentally that this is not always enough: even with a perfect world model—obtained by replacing the world model with the simulator itself, which can predict the outcome of any action sequence with zero error—a planner with a finite sampling budget still often fails to reach the goal. This can happen because planning involves two coupled operations: sampling candidate action sequences and evaluating them. A perfect world model removes prediction error from evaluation, but sampling successful action sequences can still be difficult. Motivated by this limitation, we propose IMWM (Intuition Model + World Model), which pairs the world model with an intuition model trained on demonstrations to recognize promising action sequences. Three lightweight components connect them: (i) Retrieval Initialization guides sampling by starting it from a retrieved demonstration; (ii) Hybrid Cost adds the intuition score to evaluation, alongside the world model's goal loss; (iii) Reliability Gate decides, before planning, whether to enable the intuition model. Experiments on six pixel-based goal-reaching tasks show that IMWM improves on or matches the world-model-only planner, with large gains on OGBench-Cube (93.2%, +28.5 percentage points) and Two-Room (99.3%, +10.0 points). On OGBench-Scene and DMC-Cheetah, two tasks added after all planner settings had been fixed, it gains 7.5 and 20.5 points, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.