Hierarchical W-Learning
Abstract
Inspired by a model of the brain called projective simulation, which has attracted interest within the physics community in recent years, we develop a simple and general framework for hierarchical reinforcement learning. The proposed method extends the conventional action-value Q-function to a hierarchical W-function, enabling an agent to select actions according to a hierarchy of strategies. We first present a rigorous formulation of the hierarchical framework, together with the corresponding W-learning algorithm and the hierarchical policy gradient theorem. We then demonstrate the proposed method on a navigation task to illustrate the learning procedure and its underlying mechanism. Furthermore, we extend the hierarchical framework to several modern reinforcement learning algorithms, including hierarchical deterministic policy gradient (DPG), advantage actor-critic (A2C), proximal policy optimization (PPO), and soft actor-critic (SAC), and evaluate them on the continuous navigation task along with MuJoCo control benchmarks. Experimental results show that incorporating an appropriate hierarchical strategy can consistently improve learning performance over conventional reinforcement learning methods, provided that the hierarchy is well designed and the update parameters are properly chosen.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.