acceptodds
Under review as a conference paper at ICLR 2027

Construct Before You Repair: Surrogate Contexts for Non-Prefix KV Cache Reuse

Abstract

Large language model serving increasingly reuses previously computed KV caches to avoid repeatedly prefilling long or retrieved contexts. While prefix caching is exact, more flexible reuse requires cached text chunks to be composed under different future contexts, creating a mismatch between offline cache construction and online use. Existing approaches typically encode each chunk in isolation and repair the resulting cache through online selective recomputation, effectively treating unknown future context as empty. We revisit this default and introduce SurroKV, a training-free surrogate-context cache construction method that encodes each reusable chunk behind a fixed, request-independent surrogate context and stores only the resulting chunk KV caches. The surrogate is discarded after construction and never appears at inference time, preserving both cache size and the online execution path. Our analysis shows that independent construction introduces substantial attention-routing mismatch and that surrogate conditioning can reduce this mismatch without knowing the realized future context. Across RULER, LongBench, and HELMET on multiple models, SurroKV improves accuracy by 2.1%–12.7% relative to existing methods, most notably under limited recomputation budgets and longer contexts, with no additional online latency. These results highlight cache construction itself as an important design dimension for position-independent caching.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.