BARS-Former: Robust Spiking Transformer with Bipolar Activation
Abstract
Brain-inspired spiking neural networks (SNNs) are promising for energy-efficient computing, but their performance has historically lagged behind that of ANNs. Recent spiking Transformers have substantially narrowed this accuracy gap on large-scale vision benchmarks such as ImageNet-1K. However, in this work we show that a substantial robustness gap remains under corruption and distribution shift. To probe the underlying cause, we find that vanilla spiking self-attention suffers from significantly more severe gradient attenuation on augmented samples than on clean samples, which prevents spiking Transformers from fully exploiting data augmentation for robust representation learning. To address this limitation, we propose Bipolar Activated Robust Spiking Transformer (BARS-Former), which integrates two key designs: Bipolar Accurate Spiking Self-Attention (BASA) and Ternary Pre-spike (Pre-TS). BASA adopts ternary-spike queries and real-valued keys to enlarge the effective gradient range and generate accurate bipolar attention scores, thereby improving gradient propagation from augmented samples. Pre-TS replaces the binary pre-spike activations in residual branches and downsampling modules with ternary neurons to alleviate activation truncation and gradient attenuation. Experiments on robustness benchmarks for image classification and segmentation demonstrate that BARS-Former substantially improves robustness against corruptions and distribution shifts. Further analyses show that BARS-Former effectively mitigates gradient attenuation on augmented samples, converges to flatter minima, and maintains more stable attention maps under input noise.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.