N-Spikformer: Nested Spiking Transformer
Abstract
Spiking Transformers achieve strong representation capability while retaining the energy-efficient characteristics of spiking neural networks. However, adapting them to different resource constraints typically requires repeated training or fine-tuning, which is computationally expensive. In this paper, we propose a nested spiking Transformer (N-Spikformer) that directly provides multiple sub-models with different accuracy–size trade-offs from a single network to meet different resource constraints. First, we construct a nested spiking Transformer architecture by scoring and reordering attention heads and MLP hidden neurons according to their importance. Within the nested model, smaller sub-models use the higher-ranked heads and hidden neurons. In this way, each sub-model is nested within a larger one, and all sub-models share weights. Second, we develop a nested training strategy to efficiently optimize these shared-weight sub-models. Instead of training all sub-models at each update step, we sample one representative sub-model from four representative width ratios and progressively introduce smaller sub-models during training, reducing training overhead and alleviating optimization interference. Extensive experiments across multiple spiking Transformer architectures, image classification datasets, and dense prediction tasks show that N-Spikformer directly provides customized sub-models under different resource constraints while achieving comparable or better performance than independently trained models. These results demonstrate the practical value of N-Spikformer for deployment scenarios with varying resource constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.