FastSNN: Accurate Post-Training Compression for Accelerating Spiking Neural Networks
Abstract
Spiking neural networks (SNNs) have shown promising performance on a wide range of perception tasks, making them attractive for edge deployment. However, these powerful SNNs require massive memory and computational resources, often exceeding the tight budgets of edge devices. Model compression is therefore essential, and post-training compression (PTC), which efficiently compresses a pretrained model without retraining, is particularly appealing. Existing PTC methods for SNNs minimize the local reconstruction error of each layer's signals, but this objective overlooks the discrete and temporal nature of spike transmission. In SNNs, a weight perturbation introduced by compression alters the transmitted signal only through spike flips, i.e., a spike is either inserted or deleted. We find that model performance after PTC depends less on how many flips perturbations cause than on whether flips are balanced. Building on this flip-balance perspective, we propose FastSNN, a PTC framework that jointly accounts for signal reconstruction and flip balance during SNN compression. By modeling the relation between weight perturbations and spike flips in a unified form, FastSNN applies to a broad range of compression types, including pruning and quantization. Across multiple datasets and SNN models, FastSNN outperforms the state-of-the-art method by 9.09% on average under 87.5% unstructured pruning and delivers up to a 4.2 inference speedup on a prototype neuromorphic chip.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.