acceptodds
Under review as a conference paper at ICLR 2027

SKIVE: Sparse-KV Importance-scheduled Verification by Expected true positives

Abstract

Speculative decoding (SD) drafts candidate tokens and verifies them in one target forward pass; sparse key–value (KV) cache selection reads only a budgeted subset of the cached context. Verifying the drafted sequence under a sparse cache is lossy in a way the standard acceptance count hides: sparse verification can accept tokens dense verification would reject. We score sparse verification by its true positives: the accept decisions it shares with the dense verifier. From one assumption, that a node's chance of a true positive rises linearly in the importance mass (the selector's scores) the retained blocks keep, we derive an objective for the two decisions a sparse verifier makes: which layers re-select the cache, and how sparse to go. Sparse-KV Importance-scheduled Verification by Expected true positives (SKIVE) applies it to the first, which prior work settles offline by calibrated search. Reusing one selection across all layers costs 15–27 true-positive points; the objective scores a stale selection by the mass it retains, so the best set of re-selecting layers maximizes retained mass, solved exactly by dynamic programming over a table the model builds from its own selections during decoding, and the same solve fixes the anchor count. The schedule needs no calibration data and its only model input is the layer count; on LLaMA-3.1-8B placement matters more than count: the six to eight anchors it chooses reproduce per-layer true positives, whereas six uniformly spaced layers lose 5–8 true-positive points. Against measured costs the objective becomes a throughput function Φ(sparsity, anchor count) that bounds the sparsity and says when sparse verification should, and should not, be used. End to end on LLaMA-3.1-8B with an Arctic-LSTM tree drafter at 64K, the online schedule reaches 5.5× autoregressive (AR) decoding at 4.8% lower ROUGE-L, against 3.4× for per-layer selection and 1.15× for dense SD, and within 3% ROUGE-L of AR with the same sparse cache; it transfers to Qwen3-8B with no calibration, and a fused-kernel SGLang port goes from behind dense tree verification at batch 1 to 52% ahead at batch 8.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.