acceptodds
Under review as a conference paper at ICLR 2027

When is prefix KV reuse compatible with sparse prefill?

Abstract

In sparse prefill with input-dependent attention selection, a prefix cache can depend on tokens outside that prefix. A selector may use later queries to choose the prefix's attention connections, even though the retained attention connections remain causal. We investigate when this dependence affects prefix key/value (KV) reuse in a reference implementation. Across 80 constructed document tasks, switching cache donors changed generated outputs in 11 tasks under sparse attention and 2 under dense attention. In a separate intervention, we tested 4 input pairs with identical prefixes and matched execution shapes. Dynamic routing produced different prefix KV states within each pair. These differences disappeared when the inputs shared attention connections while independently recomputing their own Q/K/V. We derive sufficient conditions for cache reuse to preserve the mathematical result. With matched deterministic execution and cache representation, these conditions also ensure that reuse and recomputation produce identical floating-point representations. On the task set, restricting reuse to complete prefill-call boundaries yielded 72 nonempty cache hits with unchanged reference outputs. Our results provide a diagnostic protocol and a conditional rule for recovering reference outputs during prefix-cache reuse.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.