acceptodds
Under review as a conference paper at ICLR 2027

PRISM: Request-Level Profitability Control for KV Cache Reuse

Abstract

Reusing key–value (KV) states is a central approach to reducing redundant computation in long-context language model serving. Existing serving systems cache shared prefixes to avoid repeated prefill, while modular caching and cache fusion extend reuse to structured context components. Cache storage, transfer, and sharing mechanisms further expand reuse across serving engines and compatible model variants. These mechanisms make candidate states available, but cache availability alone does not determine whether reuse is profitable for a request. Cache construction, transfer, and reuse preparation can consume the saved target computation, making compatibility, residency, and workload characteristics part of the latency decision. We introduce PRISM, which adds a request-level profitability admission layer after a candidate cache becomes available. The admission layer jointly checks cache compatibility, physical target residency, and a profiled cost margin, and routes rejected or failed requests to Native execution. Admitted requests execute a contiguous history suffix and the query against a read-only cache view, avoiding sparse cache assembly. On ARC and OpenBookQA, the resident executor reduces target-side latency by 37.6–53.9% in controlled long-history experiments, with the same observed accuracy as Native. On the paired evaluation, Tail is faster than the sparse-assembly proxy on 127/128 questions, with positive reductions persisting across fixed history bins, timing repeats, and tested tail ratios.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.