acceptodds
Under review as a conference paper at ICLR 2027

GRAPH-GUIDED AND ADAPTIVE KV-CACHE PREFETCHING FOR RETRIEVAL-AUGMENTED GENERATION

Abstract

The document order and meaning of the chunks required by a retrieval-augmented language model often show relationships. We have investigated whether these relationships can be used to prepare key–value (KV) cache states before a chunk is requested. We represent the chunks as a graph by including their top semantic neighbours as well as the structural links that exist within the document. For ranking the candidates, we use two-hop retrieval and select a number of chunks up to . We compare three methods: fixed weights, a globally learned balance, and an online policy which adjusts the balance after observing ranking misses. The evaluation involves an archived simulated-cache study covering six model configurations, a real experiment carried out with Qwen2.5-1.5B using vLLM and LMCache, and answer-quality experiments with Gemini 3.5 Flash-Lite. For the real access trace derived from Hotpot at , offline adaptation increases prediction recall from 80.50% to 95.33% and reduces the ready-cache time to the first token from 56.68 to 51.36 ms as compared to cosine prefetching. With synchronous candidate preparation, the corresponding times are 902.67 and 883.27 ms, whereas they are 73.41 ms when there is no prefetching. In the separate 60-question quality experiments, offline adaptation raises the HotpotQA F1 from 47.94 to 53.83, while cosine performs better on 2WikiMultiHopQA, with an F1 of 48.18 versus 40.53. Altogether, these results demonstrate the advantage of graph-guided prediction on the tested access trace and show that its impact on answer quality varies according to the dataset.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.