acceptodds
Under review as a conference paper at ICLR 2027

The Fixed Cost Is Worth Cutting: End-to-End Speedups for Activation-Sparse LLMs

Abstract

The roughly 90% of FFN intermediate rows that sparsity-trained language models leave dead at inference time cannot be skipped for free: conditional-skipping mechanisms pay costs that do not shrink with the amount they skip, and end-to-end speedups have stayed far below this byte accounting. We capture this in a compact cost model, T(a) = F + a·B/β, which yields a decision rule—cut the fixed cost. Guided by this rule, we rebuild the sparse decode path around a lean row-skipping kernel, a sparsity-aware int8 recipe, and whole-step CUDA graphs. The rebuilt stack reaches 2.66× over the model as shipped and posts 1.38×, 1.55×, and 1.94× at matched graph caliber over the official CUDA-graph baseline, a state-of-the-art sparse kernel, and the strongest measured dense stack, while sitting 1.36× above the sustained-bandwidth floor. The stack is a sparsity-trained LLaMA-2-7B under single-batch decoding on one RTX 5090, with scale probed separately (Sec. 4.4). The same experiments expose a second, temporal axis of redundancy absent in dense controls: deep FFN inputs hold a cosine of at least 0.999 across adjacent steps on 46.4% of step pairs at 7B and 82.1% at 13B, orthogonal to spatial skipping and strengthening with scale. A training-free, zero-parameter reuse mechanism cuts 18.7% off the decode step on repetition-heavy decoding (a 1.23× speedup at greedy-token agreement 1.0) at zero measured cost on natural text, with its reused steps' divergence measured and bounded. On long-chain arithmetic reasoning the full stack stays within the pre-registered tolerance with a 2.0% paired flip rate, while aggressive accelerations exceed 10%. All claims rest on a pre-registered fidelity protocol (distribution KL, decision agreement, task flip). Code for the full stack will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.