Reachability, Not Resemblance: Learning the Planning Geometry of Latent World Models
Abstract
World models (WMs) let an agent act by imagining the consequences of candidate actions. Latent WMs roll candidate action sequences forward in latent space and keep the one whose predicted final latent lies nearest the goal, so the distance behind “nearest”, the planning geometry, is what planning optimizes. Current models, LeWorldModel (LeWM) among them, use a plain Euclidean norm in an encoder trained only to predict and to avoid collapse, a geometry with no reason to reflect temporal reachability, and whether the mismatch matters has never been measured. Here we show that the mismatch is benign for near goals and dominant for far ones, and that it can be repaired without retraining: Geometry for World Models (GeoWM), a head with of the parameters fitted to the frozen encoder so that consecutive states lie within one unit and random pairs far apart, approximates temporal distance and, read as the terminal cost, restores long-horizon planning at no added planning cost. On four benchmarks the head matches or exceeds the baseline at the standard goal offset and separates from it as the goal recedes, raising success from to on Push-T and from to on Two-Room at four times that offset. The gain is in the cost, not the representation, and the head transfers unchanged to a WM of another family. The planning geometry is thus the lever for long-horizon planning, the one component prior work has held fixed while improving the rollout and the search around it. More broadly, what a planner is asked to measure matters more than the machinery around it: plan by reachability, not resemblance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.