MiST: Memory-inspired Token-Efficient Spiking Transformer
Abstract
Transformer-based Spiking Neural Networks (SNNs) combine the high performance of Transformers with the energy efficiency of SNNs, offering a promising paradigm for intelligent edge computing. However, they remain bottlenecked by the token-intensive processing inherited from conventional vision transformers, which hinders their deployment on resource-constrained edge devices. Inspired by the key-value memory mechanism of the human brain, we propose MiST, a memory-inspired token-efficient spiking transformer that exploits the inherent efficiency of transformer-based SNNs while maintaining competitive performance. Specifically, MiST decouples the full feature map into a memory repository and a small set of propagated query tokens, enabling forward inference to operate on only a fraction of the original tokens. The core component in MiST is the bidirectional spike-driven cross-attention (Bi-SDCA) module, which comprises two pathways: a forward pathway that retrieves semantic content from the repository into query tokens, and an inverse pathway that propagates enriched token representations back to update the repository. Furthermore, we observe that the memory repository in MiST exhibits high semantic disorder during early training, which slows convergence. To address this, we propose a semantic-guided self-distillation (SGSD) strategy that drives the repository to form compact semantic anchors. Extensive experiments on image classification and semantic segmentation show that MiST achieves accuracy comparable to token-intensive baselines using only 2040% of tokens, establishing a new efficiency frontier for spiking transformer architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.