Hierarchical Reinforcement Learning with Multi-Level State-Action Abstraction
Abstract
State and action abstractions can help reinforcement learning agents scale to complex environments. State abstraction groups states that can be treated equivalently, action abstraction allows agents to act with temporally extended behaviours, and combining the two at multiple levels of abstraction enables agents to learn over compact representations of an environment. Theoretical foundations exist for hierarchical state-action abstractions, but discovering such hierarchies and learning over them are open problems. We address both, proposing a temporal-difference learning architecture for multi-level state-action abstractions, with a restricted bootstrap rule that prevents value estimates at more abstract layers from biasing learning at less abstract ones. The architecture is agnostic to the discovery method, requiring only a hierarchy of nested state partitions. We construct hierarchies using graph partitioning methods. We find learning with the architecture is more sample-efficient than with the corresponding action-only abstraction in most cases, and is competitive with or better than established option discovery methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.