Q-Shaped Options for Hierarchical Reinforcement Learning
Abstract
Learning to tackle long-horizon, goal-conditioned tasks requires an agent to reason over extended timescales and act across a broad range of states. This is particularly difficult when the agent must do so entirely from offline data, without online interaction. In principle, Hierarchical Reinforcement Learning (HRL) addresses both challenges: action (temporal) abstraction reduces the effective decision horizon, while the resulting decomposition of primitive action selection into a hierarchy of decisions then permits coarser state (spatial) abstractions than in a non-hierarchical policy, as each level can discard different information irrelevant to its part of the decision process. However, realising these two benefits depends on learning an appropriate action abstraction, and current HRL methods fail in one of two ways. Some algorithms learn an action abstraction that discards distinctions between options that are needed for optimal control, undermining hierarchy altogether, while others retain unnecessary distinctions, which preserves horizon reduction, but forfeits coarser state abstraction. In this work, we characterise three desiderata for an action abstraction, and introduce Q-Shaped Options (QSO), which addresses all three. QSO leverages a prior architecture with distinct state-value functions, Q functions and policies at each level of the hierarchy to learn the action abstraction between consecutive levels as a shared encoder through their respective Q functions. The low-level Q function uses the action abstraction as a goal, encouraging it to retain distinctions between options necessary for optimal control. The high-level Q function uses the action abstraction as an action, encouraging it to discard unnecessary distinctions between options. Across offline goal-conditioned locomotion and manipulation environments, QSO learns semantically meaningful option spaces and outperforms baselines, achieving non-zero performance in tasks where all other evaluated algorithms fail. Our code is open-sourced.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.