Exploiting Temporal High Frequencies in Spiking Transformers
Abstract
Existing spiking Transformers flatten the temporal dimension into the batch axis and average over time, which is equivalent to low-pass filtering each pixel's temporal signal. We find that the temporal high-frequency components discarded by this paradigm carry discriminative information: although they account for only 18.1% of the signal energy, masking them at the input layer reduces accuracy by up to 6.2%. We further reveal a compensatory amplification mechanism: as spiking dynamics attenuate these components with depth, learnable frequency gates partially offset this decay, raising their effective utilization from 17.9% to 45.7% on CIFAR10-DVS. We propose a Spiking Spatial-Temporal Wavelet Transformer (SSTWformer), which embeds spatial-temporal Haar transforms into spiking Transformers to explicitly decompose features into eight spatial-temporal subbands for structured learning across multiple frequency bands.Experiments on six datasets spanning object recognition (CIFAR10-DVS, N-Caltech101, NCARS), gesture recognition (DVS128),and speech recognition (SHD, SSC) demonstrate that SSTWformer consistently outperforms spatial-only baselines and the attention-based Spikformer, achieving a 1.81% average accuracy gain and state-of-the-art (SOTA) on multiple tasks.The code is available in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.