acceptodds
Under review as a conference paper at ICLR 2027

Spike-BiST: Bidirectional Stabilization of Spike Representation and Surrogate Gradient for Multi-Agent Reinforcement Learning

Abstract

Spiking neural networks provide an energy efficient computing paradigm for decision making by exploiting sparse spike activity and event-driven computation. Training spiking policies under changing conditions remains challenging in multi-agent reinforcement learning. Concurrent policy updates across agents alter observation distributions and learning signals, which can destabilize spike representations and surrogate gradient propagation. To address these issues, we propose Spike-BiST, a bidirectional stabilization framework that coordinates spike representation adaptation and surrogate gradient regulation. A dual-channel adaptive population encoding mechanism is introduced to adapt receptive fields to observations and preserve consistent spike representations across training conditions. Surrogate gradient propagation is regulated by a reliability-driven modulation mechanism that combines reinforcement learning signals with spiking dynamics to reduce gradient fluctuations. Experiments on SMAC, GRF, and MPE demonstrate that Spike-BiST achieves improved task performance and training stability, with lower estimated policy-network inference energy than ANN-based models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.