acceptodds
Under review as a conference paper at ICLR 2027

STRIDE: An Efficient SNN Training-to-Deployment Framework on Real-World Vision Tasks

Abstract

Spiking Neural Networks (SNNs) have attracted increasing attention in recent years for their potential for high-efficiency and low-energy computation. The community largely agrees that realizing this potential depends on improvements in SNN training, including low-bit computing and sparse excitation, and adaptation to dedicated hardware platforms. However, evidence connecting SNN training to practical deployment benefits remains fragmented: algorithmic studies often use synaptic-operation-based analytical energy estimates, while hardware-focused studies often provide limited validation of deployed accuracy on practical tasks. In this paper, we bridge this gap with an efficient SNN TRaIning-to-DEployment (STRIDE) framework. STRIDE combines low-bit weights with Integer-LIF neurons to represent activations as compact integer codes, which are unfolded into binary temporal spikes for native, event-driven execution on dedicated SNN hardware. During training, we apply quantization-aware distillation to optimize low-bit weights under task and teacher supervision without backpropagation through time. During deployment, we develop native 1-bit spike 4-bit weight execution on a jointly developed 28-nm SNN accelerator, skipping redundant spike-weight computations when later spikes match the first-timestep reference. This work evaluates the effectiveness of STRIDE on three real-world vision tasks, achieving 61.95% mIoU on UDD6 segmentation, 31.71%/16.04% mAP/mAP on UDD detection, and 67.88% top-1 accuracy on ImageNet classification. It outperforms the strongest evaluated compact-code baselines by 3.82, 1.05/0.03, and 1.28 percentage points, respectively, while the evaluated BPTT baseline requires , , and as much peak training memory. We also evaluate hardware energy consumption by deploying one frozen STRIDE-trained SpikingLETNet checkpoint across the accelerator, GPU, and CPU. Command-block measurements yield an accelerator energy of 2.15 mJ per inference, 97.64% and 99.96% lower than on the evaluated INT8 GPU and CPU, respectively. These results also show that hardware-based energy measurements are far more accurate and reliable than conventional operation-count estimates, which capture some relative energy trends but show significant differences across architectures. Our work provides a feasible SNN implementation for real-world vision tasks and opens the door to developing the training-to-deployment collaboration scheme of SNNs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.