acceptodds
Under review as a conference paper at ICLR 2027

SSP: Self-Speculative Prefill

Abstract

Long-context inference with short outputs, common in document question answering, retrieval, and code completion, is often dominated by prefill. Existing key-value (KV) construction methods accelerate prefill by avoiding exact computation, but often make open-loop decisions without assessing whether the constructed cache entries are acceptable, creating a difficult quality-speed trade-off. This motivates identifying which draft KV entries can be safely used, rather than requiring every draft to be exact. Therefore, we introduce Self-Speculative Prefill (SSP), a framework for verification-guided cache construction with four components: construction, verification, routing, and fallback. We analyze the theoretical boundaries of KV cache construction and verification. Our realization, ACSSP, constructs KV entries using low-rank residual KV maps with original model weights frozen, verifies drafts with a training-free attention-weighted verification score, uses token-level routing to select aggressive acceptance or continuation to a conservative anchor, and reuses verification-boundary states during continuation. Compared with POP, ACSSP improves average quality on LongBench and RepoBench by up to 7.47 points, delivers up to and TTFT speedups on RULER at 32K and 128K, respectively, and reduces mean prefill matmul FLOPs by up to 20.19% across seven downstream tasks on Llama-3.1-8B-Instruct. Code is available at https://anonymous.4open.science/r/Zgrnf7sg.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.