ACT-SNN: Activity-Constrained Two-Stage Conversion for Efficient Spiking Large Language Models
Abstract
Large language models (LLMs) incur substantial computational and energy costs during inference, motivating their conversion into energy-efficient spiking neural networks (SNNs). Although existing LLM-to-SNN methods have achieved high conversion fidelity, this does not necessarily guarantee efficient inference: accurate discrete representations can require high spike activity and thus higher energy consumption, while layer-wise temporal accumulation leads to serial inference latency. In the first stage, our Basis-Augmented Multi-Threshold Neuron (BA-MTN) jointly learns basis transformations and channel-wise thresholds under an explicit activity constraint, decoupling reconstruction optimization from spike activity control while enlarging the feasible space for a favorable reconstruction–activity trade-off. In the second stage, the Basis-Augmented Symmetric Integrate-and-Fire (BA-SIF) neuron provides an equivalent temporal realization with causal step-by-step spike propagation, enabling pipelined cross-layer execution without layer-wise temporal accumulation; relaxation steps and differential operators further compensate for temporal propagation delays and nonlinear operations. Experiments across LLaMA-2, LLaMA-3, and Qwen2.5 models ranging from 7B to 70B parameters demonstrate that ACT-SNN achieves a strong accuracy–energy trade-off under both full-precision and 4-bit weight settings. In particular, under 4-bit weights, ACT-SNN attains energy ratios of –, reducing estimated inference energy by –% over representative W4A4 QANN baselines while maintaining competitive zero-shot accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.