Planning with Value-Aware Learned Goals for Hierarchical Reinforcement Learning
Abstract
Hierarchical reinforcement learning (HRL) decomposes long-horizon decision making across temporal scales, with a high-level policy periodically specifying goals and a low-level controller realizing them through primitive actions. Departing from the conventional use of learned goal-conditioned policies for low-level control, we introduce Value-Aware Latent-Space Planning (VALP), a novel low-level policy-free framework that combines reinforcement learning for strategic decision making with world model-based planning for closed-loop control. VALP learns a compact, value-aware latent space that preserves task-relevant information, in which a high-level policy specifies temporally extended abstract goals, while a model-predictive planner directly synthesizes primitive actions by planning through a learned latent world model toward these objectives. By coupling value-based representation learning and model-based planning, VALP separates learning task-relevant objectives for long-horizon decision making from their realization through short-horizon control. VALP significantly outperforms state-of-the-art HRL methods on long-horizon control tasks. Our findings highlight the importance of a task-relevant communication interface for hierarchical control and demonstrate how reinforcement learning for long-horizon decision making can be effectively combined with world model-based control for goal realization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.