It's the Model, Not the Macro-Actions: What Hierarchy Buys in Latent-Space Planning
Abstract
A hierarchical latent world model raises two abstraction levels at once: it rolls plans through a model that predicts ten steps at a time, and it searches over learned temporally extended macro-actions. The pair beats flat planners on long-horizon tasks, and the standard explanation credits the action abstraction. Using a released model pair whose flat and hierarchical checkpoints have bitwise-identical level-1 weights — so arms differ only in the planner and in which of the two predictors rolls the plan — we find that, on this pair, the credit belongs elsewhere. We first characterize what the learned macro-action space contains: a -dimensional code whose variance the two lowest temporal modes of a ten-step action chunk reach to and plain constant-hold chunking to . We then run the same hand-designed action classes through each of the two models. Over the fine model, every restriction scores below the unrestricted planner in the same harness — , , , across classes spanning to coverage, against (three of the four gaps significant). Over the coarse model, the crudest class — two numbers per macro-action — reaches , matching the full released hierarchy, with 14 discordant trials to 0 against the identical class over the fine model (); searching the entire learned -D space does no better (). Model-side measurements taken with no planner in the loop suggest why: at long lookahead the fine model's cost stops ordering plans the way the environment does (latent cost, from to steps; position cost, at ), while the coarse model's does not, and the gap is significant at steps (, , 17/20 layouts, plans). At steps an earlier estimate of that gap (, ) did not survive a larger plan sample (, ); we report it as within noise. We also report what failed: truncating the flat planner's lookahead to where its model is still trustworthy does not help, monotonically, across a range (, , , ), because the task's own objective is uninformative at short horizons ( at 50 steps). Of this paper's preregistered predictions three were confirmed, three refuted and two null: the confirmations are two model-side measurements and the prediction that constant hold over the coarse model reaches ; every refutation is a prediction about the flat planner at its own timescale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.