HRP: Hierarchical Residual Prediction for Latent World Models
Abstract
A world model is only useful for long-horizon planning if the agent can search it, and search is only affordable if most of the search can be skipped. One that predicts every step the same way gives no indication of which steps matter, so a hierarchy built on it runs its coarse level regardless of need. We introduce Hierarchical Residual Prediction, a latent world model whose upper level predicts the lower level's error instead of the future. The correction it produces is near zero wherever the lower level is already accurate, and because it is a vector rather than a score it carries how the lower level failed and not only that it failed. Dynamics the agent does not control produce corrections as large as a decision point does, so the size of the correction identifies decision points no better than chance, while the way it varies with the action identifies them reliably. Decision points therefore emerge from offline, reward-free transitions with no labels, rewards, or tuned sparsity, and restricting coarse search to them improves long-horizon success while reaching the best success of prior hierarchies at 2.4x fewer model rollout steps. The construction recurses, and a third level commits at timescales the second cannot reach.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.