acceptodds
Under review as a conference paper at ICLR 2027

QSpar: Co-Designing Sparse Attention and Quantized Weights for Long-Context Self-Speculative Decoding

Abstract

Reducing KV-cache traffic alone is insufficient to eliminate drafting overhead in long-context self-speculative decoding. Under a fixed GPU memory budget, longer contexts constrain the feasible batch size, limiting the amortization of weight access across requests. A sparse-attention-based drafter further reduces draft-side KV-cache traffic but leaves repeated access to the target-sized drafter weights unchanged, exposing a substantial weight-side bottleneck. Low-bit weight quantization reduces this cost, but naively combining it with sparse attention degrades acceptance length. We address this runtime–acceptance trade-off through two key observations: token flips concentrate at low-margin positions, and the BF16 target token usually remains among a few top-ranked draft candidates. Based on these observations, QSpar combines a weight-quantized sparse drafter with a candidate Reranker and a margin-aware Reducer. The Reranker recomputes candidate logits using BF16 LM-head weights, while the Reducer selectively retains alternatives under a bounded candidate budget. Final verification uses the original full-attention BF16 model to preserve the exact target-model generation. Under a fixed KV-cache budget, QSpar achieves average and maximum end-to-end speedups over Vegas, and up to over vLLM.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.