acceptodds
Under review as a conference paper at ICLR 2027

Decomposed Contrastive Critics for Continual Reinforcement Learning

Abstract

Continual reinforcement learning is commonly studied with dense, task-specific rewards, although many control problems are more naturally specified by goals and sparse feedback on task completion. We study this setting on a sequence of goal-conditioned manipulation tasks and ask how contrastive reinforcement learning can transfer knowledge across task boundaries. Because contrastive critics must capture both reusable goal-reaching structure and task-specific action preferences, transfer requires the former to persist across tasks while the latter remains adaptable. We introduce Decomposed Contrastive Critics (DCC), an algorithm that decomposes the critic into a persistent shared state-action representation and a task-specific representation re-initialized at each task boundary, coupled with a success replay buffer and an action-sensitive critic gate that deploys targeted policy supervision when the critic cannot differentiate actions. On the ten-task Continual World V2 Sawyer sequence, we compare against nine contrastive transfer baselines and two sparse goal-conditioned SAC baselines. DCC improves both continual adaptation and task success by a substantial margin, successfully solving all ten manipulation tasks in the curriculum.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.