Bridging Efficiency–Accuracy Gap in Spiking Neural Networks via Operation-Constrained Quantization and Low-Spike Suppression
Abstract
Spiking neural networks constitute a promising approach to energy-efficient intelligence by leveraging binary, event-driven computation. However, binary spike signals substantially constrain the channel capacity for information delivery, thereby limiting attainable model accuracy. To jointly improve representational capacity and inherit event-driven computation, we propose a novel multi-level spiking neuron termed QuantLIF, consisting of operation-constrained quantization and low-spike suppression. Specifically, we construct a dense quantization codebook whose values can be represented by a bounded number of signed power-of-two terms. Through shift-based re-parameterization, synaptic integration is replaced by several accumulations, without introducing excessive inference delay. To counteract the deterioration in event sparsity caused by multi-level spikes, the low-spike suppression stochastically filters low-contribution spikes during training and deterministically eliminates them during inference. We theoretically analyze QuantLIF from the perspectives of quantization and information theory, respectively. Comprehensive experiments demonstrate that QuantLIF achieves a favorable trade-off among predictive accuracy, firing sparsity, and energy consumption across both static and neuromorphic benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.