Spiking Actor-Critic with Temporal Dendritic Heterogeneity for Asynchronous Multi-Agent Reinforcement Learning
Abstract
Asynchronous multi-agent reinforcement learning requires agents to coordinate while making decisions at different macro-action completion times. Information exchange can supplement incomplete local observations, and yet under asynchronous execution, a receiver's current observation can be combined with teammate messages generated at the same or earlier times. This temporal mismatch can make it difficult to retain useful teammate information without allowing outdated information to influence current decisions. To integrate this asynchronous information, we propose CoSpAC, a compartmentalized spiking actor-critic framework inspired by temporal dendritic heterogeneity. Each actor processes local observations and aggregated teammate messages through separate dendritic compartments with distinct decay factors. In this way, CoSpAC provides different temporal integration timescales for observations and messages from different decision times, which are subsequently combined through somatic dynamics for action selection. Additionally, CoSpAC exchanges teammate information at agents' individual decision events and optimizes message generation through a receiver-oriented objective based on its influence on receiving agents' policies. Experiments across asynchronous multi-agent tasks show that CoSpAC outperforms state-of-the-art asynchronous MARL approaches, demonstrating the effectiveness of the proposed framework for asynchronous cooperation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.