GASPOL : Dual Gating Sub-Networks for Continual Multi-Agent Reinforcement Learning
Abstract
Continuous coordination under dynamic environments remains a challenge for a cooperative multi-agent system. The existing study on Continual Multi-agent Reinforcement Learning (MARL) focuses on task discrimination that requires trajectory buffers from previous tasks, which could be inapplicable in real-world applications. In addition, the existing study requires feedback or numerous trajectory data to select the best available action layer during the inference phase. Addressing the mentioned issues, this paper proposes a study on rehearsal-free continual multi-agent reinforcement learning and proposes a new method named Dual Gating Sub-networks Policy (GASPOL). GASPOL introduces a dual expert system, i.e., global experts and local experts, as a part of the trained policy network. Global experts preserve general knowledge (stability) of all learned tasks, while local experts obtain task-specific knowledge (plasticity). GASPOL leverages a new actor advantage-guided regularization mechanism for more flexible regularization at each training iteration. Our evaluation on MARL environments, i.e., SMAC, Pettingzoo MPE2, and MAMuJoCo shows that the proposed method significantly achieves better stability-plasticity than the existing state-of-the-art methods. Our extended analysis shows that the proposed method can maintain its performance under various numbers of experts and neurons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.