Asymmetric Membrane Credit for Spiking Actor–Critic Learning
Abstract
A stateful spiking core produces both binary events and continuous membrane states, but an event-only Critic does not directly read the latter. We therefore introduce Asymmetric Membrane Credit (AMC), a training-time interface that gives the Critic a zero-initialized readout of the shared pre-threshold membrane without adding an Actor inference branch. A matched detached control retains the trainable readout while blocking only its additional encoder-gradient route, separating membrane-informed Critic learning from direct encoder credit. A reusable stateful spiking Actor–Critic substrate implements AMC across PPO, SAC and TD3, with common policy computation and rule-specific state reconstruction and encoder updates. Across six MuJoCo-v5 tasks and three paired training seeds per setting, the complete AMC pathway achieves higher mean area under the learning curve (AUC) than event-only controls in 17 of 18 settings. Fixed final policies achieve higher mean return than the same controls in 16 settings when evaluated on 20 common reset seeds. Direct credit improves both AUC and final-policy return on SAC HalfCheetah but reduces both on SAC Walker in all three training pairs. In a separate Ant-v4 system comparison at the same interaction budget, Membrane-Credit PPO (MCrPPO) improves mean AUC by 64.8% over a SpikeGym reproduction retaining its native training schedule. Paired GPU and neuromorphic evaluation verifies Actor-only execution of selected MCrPPO policies after Critic removal. Together, AMC and the substrate combine improved training performance with explicit control over state access and encoder updates, while retaining an Actor-only deployment path.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.