An Expert Knowledge Composition Method for Continual Reinforcement Learning
Abstract
Continual reinforcement learning requires an agent to continuously adapt to sequentially arriving tasks while retaining and reusing previously acquired knowledge. However, when past interaction data are no longer accessible, effectively organizing historical knowledge, identifying transferable knowledge components, and preventing the knowledge repository from growing unboundedly with the number of tasks remain key challenges. To address these issues, we propose an expert knowledge composition method for continual reinforcement learning (EKC-CRL). The proposed method represents each historical task as a group of expert knowledge vectors, enabling a more fine-grained characterization of task-specific knowledge than a single task-level representation. When adapting to a new task, the model performs two-level expert routing to select and compose relevant historical experts, while maintaining an independent set of current experts to acquire task-specific new knowledge. We further introduce routing-balancing and expert-diversity regularization to improve expert utilization and reduce knowledge redundancy across experts. As tasks arrive sequentially, learned experts are incorporated into a historical knowledge pool. Once the pool reaches its capacity limit, Hungarian matching is employed to align experts across tasks, followed by the merging of similar knowledge groups. In this way, the model can achieve flexible reuse of historical knowledge and effective new-task learning without storing past interaction trajectories. Experimental results demonstrate that our method consistently outperforms state-of-the-art approaches across multiple continual reinforcement learning benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.