Are Scaling Laws Universal in Spiking Large Language Models?
Abstract
The scaling law is a foundational property of large language models (LLMs), dictating predictable performance scaling in foundation models. Spiking LLMs are often presented as neuromorphic analogues of quantized models which convert continuous-valued LLM activations into discrete temporal spike trains. However, whether Spiking LLMs genuinely exhibit consistent scaling laws remains an open and unresolved question. In this paper, we demonstrate that the validity of scaling laws in Spiking LLMs depends fundamentally on whether the system is analyzed under a static conversion regime or a temporal dynamical regime. Under the static accumulate-then-fire (ATF) regime, where synaptic input is fully accumulated prior to firing, conversion acts effectively as a static quantizer. The excess loss exhibits a predictable representation-limited scaling law governed by network scale and spike-code capacity . In contrast, when operating in the temporal dynamical regime via online charge-and-fire (OCF), which represents the asynchronous parallel mode native to neuromorphic hardware, scaling laws become strictly conditional. In this dynamical regime, small time windows and coarse threshold levels let temporal unevenness dominate, causing excess loss to deviate from power-law scaling against or . Power-law scaling is only restored when and code capacity exceed critical dynamic thresholds. Our findings reveal that scaling claims for SLLMs cannot treat them simply as bio-inspired quantized LLMs, but must explicitly specify the dynamical regime and the condition under OCF.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.