No Passage Is an Island: Orphaned KV Caches, Phantom Peaks, and Warm Encoding for Retrieval-Augmented Generation
Abstract
Retrieval-Augmented Generation improves Large Language Models by enabling document-grounded generation, however it requires recomputing the per-passage prefill at each retrieval time. As such, the systems cache each passage's KV states once and assemble the retrieved ones at query time. Current methods reduce this overhead by reusing the per-passage KV cache of prior prefills, so that generation can skip the prefill of retrieved passages. The assembled cache, however, reduces prediction quality, and prior systems repair it online by recomputing roughly a quarter of the tokens on every query. In this work, we take a different approach. We identify where the error comes from and remove most of it offline. We find that position is not the problem, but rather that a passage encoded alone is encoded cold, meaning that its keys are wrong in the low-frequency rotary bands, exactly where the question attends. The question's attention therefore collapses and lands on phantom peaks, tokens that the mismatched question favours and the fresh prefill does not. We introduce Raftel, which works by giving the phantom peaks their true keys, so the tokens to recompute are the ones the mismatched question attends, not the ones the true question would. We therefore warm the passages instead of repairing the cache. Each passage is encoded once against the stored caches of its most similar passages and iterated twice, which removes most phantom peaks before any question arrives. We evaluate our method on LongBench and RULER and find that the bare warm cache matches the prior state of the art's best operating point at 6.6x the speed of full prefill.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.