Continual Reinforcement Learning with Quasimetric Structure Evolving across Tasks
Abstract
The goal of continual reinforcement learning is to integrate accumulated experience, rapidly learn new tasks, and minimize forgetting. In real-world scenarios, transition dynamics and reward functions often vary substantially across tasks, rewards are frequently sparse, and abrupt changes occur at task boundaries, making it difficult to learn a task-agnostic policy that generalizes across tasks. Inspired by compositional recombination in human cognition, we propose CONQUEST, a framework that learns a cross-task quasimetric representation to capture asymmetric and compositional knowledge-transfer relationships during continual learning. Quasimetric structures naturally characterize the reachability between states and goals, enabling agents to identify transferable knowledge without relying on task-specific rewards and transition dynamics. We introduce the cross-task compositional successor distance that connects task-specific representations through shared latent waypoints, and further learn a unified quasimetric space to capture reusable structures across tasks. CONQUEST enables effective task-agnostic policy extraction, facilitating the reuse of prior experience while mitigating forgetting as the task sequence expands. Our approach outperforms existing baselines on continuous-control tasks under both dense- and sparse-reward settings, as well as on image-based continual reinforcement learning benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.