Deferred Contextualization for Efficient Long-Context KV Reuse
Abstract
Shared-corpus long-context QA repeatedly queries the same document blocks. Per-query re-encoding is expensive and several cache reuse mechanisms have been proposed recently offering varying levels of quality-speed tradeoffs. We analyze this tradeoff through C4: four factors that govern whether reusable caches preserve multi-block reasoning quality: Compactness of relevant blocks relative to the instruction, Contiguity among them, Cross-attention across them, and Cleanliness, requiring each relevant block to be encoded without attending to unrelated predecessors. Through controlled experiments we show that each factor matters: accuracy drops when relevant blocks are far from the instruction or from one another; independently cached blocks lose cross-block interaction; and encodings contaminated by irrelevant blocks cannot be fully recovered by decode-time masking alone. Motivated by the C4 lens, we propose PRISM, a training-free KV reuse pipeline. PRISM encodes each block independently alongside the instruction, yielding clean reusable KV states, and reassigns relevant blocks to compact, contiguous positions at query time. PRISM establishes cross-block dependencies during query time through two operations: iterative span extraction, which conditions each block's extraction on the query and on spans extracted from preceding blocks; and in-place KV recomputation, which updates only the extracted span states with full cross-attention. Across long-context QA, claim verification, and synthetic graph traversal tasks, PRISM recovers much of the quality of per-query re-encoding while reducing time-to-first-token by nearly 5×, without fine-tuning or full re-encoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.