Remembering How to Act through Greedy Action Consolidation in Continual Reinforcement Learning
Abstract
For value-based continual control, retaining a learned decision need not entail preserving the action values that originally supported it, since greedy action selection depends on their relative ordering rather than their absolute magnitudes. We propose Greedy Action Consolidation, or GAC, which targets the gaps that support historical greedy actions to mitigate catastrophic forgetting while allowing action value estimates to evolve. At the core of GAC lies the continual action gap, which measures the action value difference, under the updated function, between a historical greedy action and its best competing action, and is regularized to remain at least a prescribed fraction of it's historical action gap. Theoretically, we derive an upper bound on return discrepancy in terms of the GAC regularizer and the distribution of historical action gaps, and show that GAC strictly enlarges the pointwise admissible action value set relative to exact value preservation as instantiated by Value Cloning, affording greater freedom for value adaptation. Experiments on MinAtar and Forager demonstrate the effectiveness of GAC across semi-continual and continual reinforcement learning settings, suggesting that the action gaps provide a parsimonious substrate for remembering how to act in continual reinforcement learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.