acceptodds
Under review as a conference paper at ICLR 2027

Spec-Spike: Spiking Drafters for Vision-Language-Action Speculative Decoding

Abstract

Spiking neural networks (SNNs), the native substrate of neuromorphic hardware, compute with sparse spike events and need far less synaptic arithmetic than dense networks. Speculative decoding of vision-language-action (VLA) policies relies on dense drafters whose action-token proposals a 7B-parameter verifier checks in one pass, and a spiking drafter should make these proposals far cheaper. Making only the drafter's core spiking, however, cuts the core's cost over 15x but the drafter's only about 3x, because the dense projections that feed the core and read out its logits recur at every draft position and take 87% of its analytical cost. We therefore design Spec-Spike around these interfaces in two coupled steps. First, it restructures each interface by how often its input changes: context is cached per observation, token projections are tabulated, and the hidden-state and context projections are factored to low rank. Second, because this factorization costs a spiking drafter 14% of its accepted prefix, against 5% for a dense one, and more than halves its spike activity, Spec-Spike recovers the factors on its own prediction histories. On LIBERO-Long, Spec-Spike reaches mean task success not detectably different from a parameter-matched Spec-VLA-style dense drafter with the same interface optimizations (49.93% vs. 48.21%) at 11.35x lower analytical drafter cost, and a second branch, still 6.35x cheaper in drafter arithmetic than one dense branch, accepts longer prefixes than the single-branch dense drafter with fewer verifier passes. The released Spec-VLA drafter, run with official Spec-VLA and KERV code, spends 750-1,540x more drafter arithmetic.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.