acceptodds
Under review as a conference paper at ICLR 2027

Not Forgotten, Not Retrieved: A Same-Checkpoint Decomposition of Failure in Modular Continual RL

Abstract

An old-task performance drop in continual reinforcement learning can mean that a learned capability degraded or that the agent failed to select a capability that remains available. We introduce an architecture-agnostic diagnostic protocol for modular agents whose expert selection can be overridden, and instantiate it in CM-Dreamer, a modular DreamerV3 agent. At the same checkpoint, we compare a task aware reference that forces the expert assigned to the evaluated task with learned deployment, in which the agent’s router selects or mixes experts without a task label. Their signed difference is the retrieval gap; change in task-aware reference performance across training is retention loss. Acquisition is measured separately. In CM-Dreamer, the same final checkpoint on a three-task Meta-World stream (cw3) scores 0.66 under the task-aware reference and 0.32 under learned deployment. Another checkpoint places its largest router weight on the assigned expert on 97% of steps yet has a retrieval gap of 1.00, whereas inaccurate selection can be harmless when several experts solve the same task. Forced experts and fixed mixtures distinguish these cases. When observations reveal task identity, an observation-based selector without task labels matches the reference. When identity-revealing features are masked, a tested router identifies the task only after early actions make correction ineffective. Freezing the reference computation preserves old-task performance but often prevents new-task acquisition; unfreezing the recurrent state model and decoder improves acquisition at a retention cost. Most checkpoints contain only one usable skill because later tasks were often never acquired. The few two-skill checkpoints demonstrate selection failure but are too few to establish its frequency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.