SSP: Self-Speculative Prefill
Abstract
Long-context inference with short outputs, common in document question answering, retrieval, and code completion, is often dominated by prefill. Existing key-value (KV) construction methods accelerate prefill by avoiding exact computation, but often make open-loop decisions without assessing whether the constructed cache entries are acceptable, creating a difficult quality-speed trade-off. This motivates identifying which draft KV entries can be safely used, rather than requiring every draft to be exact. Therefore, we introduce Self-Speculative Prefill (SSP), a framework for verification-guided cache construction with four components: construction, verification, routing, and fallback. We analyze the theoretical boundaries of KV cache construction and verification. Our realization, ACSSP, constructs KV entries using low-rank residual KV maps with original model weights frozen, verifies drafts with a training-free attention-weighted verification score, uses token-level routing to select aggressive acceptance or continuation to a conservative anchor, and reuses verification-boundary states during continuation. Compared with POP, ACSSP improves average quality on LongBench and RepoBench by up to 7.47 points, delivers up to and TTFT speedups on RULER at 32K and 128K, respectively, and reduces mean prefill matmul FLOPs by up to 20.19% across seven downstream tasks on Llama-3.1-8B-Instruct. Code is available at https://anonymous.4open.science/r/Zgrnf7sg.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.