QueryLens: Lightweight Query Recalibration for Reused KV Caches
Abstract
Long-context language model serving increasingly relies on key–value (KV) cache reuse to avoid repeated prefill computation over recurring documents. Yet caches compiled independently and assembled under a new request cannot fully reproduce the contextual interactions of joint prefill and require correction. Existing approaches typically operate on the cache side through selective token recomputation, aiming to make reused caches more consistent with their full-context counterparts. However, cache-side repair can increase time to first token (TTFT) as the recomputation budget grows, while the native query may remain mismatched with the realized cache. To address this issue, we propose **QueryLens**, a query-correction method that introduces a new query-side perspective on KV cache reuse. It learns a lightweight, shared transformation of queries to improve how reused KV states are read. Moreover, QueryLens supports token-level and attention-head-level calibration and can be combined with existing KV cache methods to further improve performance; its statically folded head-level variant incurs negligible measured TTFT overhead. Experiments show that adding QueryLens to A3 at recomputation raises the three-task macro-F1 on Llama-3.1-8B-Instruct from to .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.