acceptodds
Under review as a conference paper at ICLR 2027

Async-SNN: Efficient Spike-Driven Large Language Models with Asynchronous Computing Paradigm

Abstract

Large language models (LLMs) are increasingly deployed in edge scenarios where power supply can be strict, intermittent, and time-varying. Spiking neural networks (SNNs), with their event-driven and asynchronous computation paradigm, offer a promising path toward power-adaptive LLM inference. However, existing spiking LLMs largely inherit synchronization-intensive operators from conventional Transformers, such as RMSNorm and Softmax Attention, whose global statistical dependencies fundamentally prevent true neuron-level asynchronous execution. In this work, we propose Async-SNN, a framework for constructing fully spiking LLMs with asynchronous computing capability at the billion-parameter scale. We first build a ready-to-convert Async-LLM by replacing RMSNorm with Hardtanh Dynamic Normalization (HDN) and Softmax Attention with Normalized Linear Attention (NLA), thereby eliminating major sources of synchronization. We then introduce a coarse-to-fine ANN-to-SNN conversion pipeline that combines calibration, supervised fine-tuning based quantization-aware training, and equivalent ST-BIF neuron conversion to preserve task performance. Finally, we propose greedy adaptive firing (GAF), which dynamically reallocates neuron firing budgets across layers and time-steps under instantaneous power constraints. Experiments on 0.6B and 1.7B models show that Async-SNN consistently outperforms existing spiking LLM baselines and maintains strong perplexity and zero-shot accuracy even when only 10% of neurons are allowed to be active per time-step.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.