HOMER: Rational Meta-Reasoning for Hierarchical Reinforcement Learning
Abstract
A key capability for sequential decision-making agents is *rational meta-reasoning*: the ability to adaptively balance the requisite computing time to solve a problem against the time taken to perform those computations. Rational meta-reasoning agents can compute for the minimum amount of time needed to generate high-quality decisions, thereby avoiding poor decisions from insufficient computation or wasteful computations with marginal utility. In this work, we study meta-reasoning in the context of model-free hierarchical reinforcement learning (HRL) agents, composed of a manager who assigns -step subgoals to a worker. Since each level in the hierarchy optimizes a different objective, at a different temporal horizon, it raises the question of whether the requisite computation at each level is different? Motivated by this question, we first formalize meta-reasoning for HRL agents, and then introduce HOMER a hierarchical rational meta-reasoning agent, that instantiates the meta-reasoning process through iterative computations within model-free recurrent neural networks. We then apply HOMER to the offline goal-conditioned RL setting and demonstrate that it successfully attains goal-reaching success rates on par with non-meta-reasoning agents while using far less computation. We then empirically study the adaptive computations deployed at each level of the hierarchy; we observe that the worker's adaptation of compute correlates with distance to subgoals whereas the manager's adaptation of compute correlates with the entropy of the empirical -step distribution induced by the offline dataset. Finally, we show that a meta-reasoning manager can cache previous computations to improve its goal-reaching success rates and meta-reasoning efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.