Continual Memory: Separating Selection and Representation in Continual Reinforcement Learning
Abstract
Rehearsal reduces forgetting in continual reinforcement learning by storing a small memory of past states and constraining the policy on them. Yet this memory is typically formed by uniform sampling, leaving open which states should be retained and how they should be compared. We introduce Continual Memory, a framework for rehearsal memory construction that separates two decisions: a representation that defines similarity between states and a selection rule that determines which states are stored. On Continual World, we vary these decisions separately while holding the policy, training budget, and rehearsal loss fixed. The selection rule changes retention. Coverage based selection increases forgetting in every representation we test, whereas density based selection shows no detectable increase in forgetting and forgets less than coverage on the same seeds. With the selection rule fixed, the representation changes forward transfer. A merge and reduce representation of policy gradients reaches forward transfer, compared with for our ClonEx-SAC reproduction and reported previously, although this improvement is supported by the unpaired analysis and comes with increased forgetting. Continual Memory with density selection reaches forward transfer while preserving ClonEx-SAC's performance and retention. These results separate two roles of rehearsal memory construction: selection determines retention in the fixed representation we study, while representation determines forward transfer under coverage selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.