SpikeHENA: Spiking Hybrid Encoding and Nonlinear Approximation for Training-Free LLM-to-SNN Conversion
Abstract
Spiking neural networks (SNNs) are emerging as a promising alternative to conventional Transformer-based large language models (LLMs) because of their spike-driven computation and sparse communication. However, existing ANN-to-SNN conversion methods face a difficult trade-off among representation fidelity, spike activity, and the required number of time steps, while key nonlinear operations such as Softmax, RMSNorm, and SiLU remain difficult to realize with signed binary spikes. These limitations hinder high-fidelity conversion of pretrained LLMs. In this work, we address the conversion problem from both spike encoding and nonlinear computation perspectives. For spike encoding, we propose a Two-Stage Hybrid Encoding Neuron that maintains a single membrane state across both stages, using population-activity-constrained BinarySpike encoding for large-magnitude activation components and SingleSpike encoding for the remaining residuals. For nonlinear computation, we design dedicated spiking units for square root, exponentiation, and division, and compose them into spike-driven implementations of Softmax, RMSNorm, and SiLU. Based on these designs, we develop SpikeHENA, a training-free conversion framework that directly reuses pretrained ANN parameters. Experiments on BERT and four LLMs show that, at , the converted LLMs incur only 0.12–0.64 percentage-point drops in average zero-shot accuracy and small perplexity increases on WikiText-2 and C4, while their theoretical synaptic-computation energy is estimated at 31.2%–32.5% of the corresponding ANN estimates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.