SPECULATIVE CACHE LINE PUNCTURING: GRANULAR MEMORY REDUCTION FOR LLM INFERENCE
Abstract
Long-context speculative inference frequently suffers from severe memory bottlenecks caused by rapidly expanding Key-Value (KV) cache footprints. Existing eviction and quantization strategies often introduce intra-cache fragmentation or disrupt token-index alignment between draft and target models. In this paper, we introduce Speculative Cache Line Puncturing (SCLP), a fine-grained runtime memory-management primitive tailored for speculative execution. Rather than evicting coarse blocks or reallocating memory, SCLP applies an in-place sub-block validity mask to deactivate redundant attention states while strictly preserving underlying tensor layouts and positional coordinate systems. We evaluate SCLP on the Qwen-2.5 architecture (7B target with 1.5B draft) across three distinct hardware tiers (NVIDIA A40, A100 SXM 80GB, and H100 SXM 80GB) using 500-sample evaluations across 5 independent random seeds (2,500 cumulative evaluation prompts per benchmark). Our findings reveal a clear operational dichotomy: on short-context conversational flows (ShareGPT), SCLP maintains low kernel overhead and yields an acceptance rate improvement of 53.2% vs. 42.9% (+10.3%, ) while leaving short-context memory footprint unperturbed; conversely, on long-context retrieval workloads (RULER, 4k–16k tokens), SCLP achieves a sustained 48.2% reduction in active KV cache memory (0.33 GB vs. 0.65 GB per sequence at 4k context) and delivers a +12.7% boost in speculative acceptance rate (61.5% vs. 48.8%, ), nearly doubling concurrent serving batch capacity under fixed VRAM constraints.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.