LETHE: Forgetting Irrelevant Context for Efficient Prefill in RAG and Tool-Calling
Abstract
Agentic, retrieval-augmented generation (RAG), and tool-calling applications repeatedly process largely unchanged context, yet standard KV caching only applies when the reused context forms an identical prefix. We introduce LETHE, a training-free method for non-prefix KV reuse. Rather than caching chunks independently and later reconstructing missing cross-attention, LETHE builds a shared cache in which chunks already interact, then selectively corrects the influence of chunks excluded from the current request. An offline divergence analysis is performed once per chunk to identify tokens that should be recomputed at runtime, while the remaining tokens are reused from cache. The same mechanism applies to both retrieved text and tool definitions without modifying or fine-tuning the base model. Across Qwen3-8B, Qwen3-14B, and Llama-3.1-8B-Instruct, LETHE recomputes as little as 4.9% of context tokens and reduces full-prompt FLOPs by 50.7–87.5%. On HotpotQA, the best evaluated LETHE setting for each model is within 2.8 EM points of full recomputation and LETHE-K outperforms EPIC-16 by 7.1 EM points on average; on BFCL, the best LETHE setting is within 1.5 percentage points of full-recomputation function-call accuracy across all three models. With a fixed 10% recomputation budget, LETHE achieves 5.3–6.3× wall-clock prefill speedup at 10k context tokens on a single RTX A6000 GPU. Overall, across three models and both text retrieval and tool calling, LETHE preserves near-baseline task quality while substantially reducing computation, with long-context prefill speedups reaching about 6×.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.