How Much of a Policy Library Must Be Kept in Continual Offline RL?
Abstract
Continual offline reinforcement learning trains a specialist policy for each arriving task from static datasets, and a natural way to avoid forgetting is to keep every specialist in a growing library. Yet tasks share behavior, so how much of this library must be retained to keep serving every task seen so far? Task labels and parameter isolation cannot answer this question, since neither reveals whether an existing policy can actually solve a task. This paper answers this question with behavioral evidence. Paired-reset validation tests each candidate policy directly on each task identity and builds a directed compatibility graph. From this graph, we select a cover that decides both which policies to keep and which task each kept policy serves. The same behavioral evidence also bounds its own cost. Once every completion of the unobserved edges yields the same retention decision, further validation cannot change the outcome, and we can skip it. On 3D navigation benchmarks, our framework covers every task identity while retaining a minority of the specialists, improves service success over per-identity retention, and removes a substantial fraction of candidate validation episodes without altering any retention decision. It thus provides a simple, verifiable answer to what must be kept, and to when further validation is no longer necessary.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.