Online Policy Optimization for Autonomous Driving with a Hierarchical World Model
Abstract
Policies obtained via imitation learning are known to struggle with out-of-distribution cases, i.e., when the environment or the agent deviates from the expert distribution at test time. Test-time planning offers a potential solution by learning a more generic world model and reward model from demonstrations and searching for corrective actions at inference. However, this approach has not been verified in driving setups due to the computational cost of searching over long planning horizons. We introduce HiPlan, a hierarchical planning framework that mirrors the temporal structure of the hierarchical policy it supports, with an upper level operating at the subgoal timescale and a lower level at full resolution. This enables two-stage search, which reduces search dimensionality by an order of magnitude. We find that the two levels develop qualitatively different representations, with the upper model encoding strategic features and the lower encoding tactical features, a double dissociation confirmed by linear probing and horizon-scaling analysis. HiPlan achieves state-of-the-art on nuPlan and CARLA, and analytical experiments show that the hierarchy enables more efficient search, that representations factorize by timescale, and that structural alignment between policy and world model is necessary for both.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.