Reusing Retained History for Fast Prefill after Agent Context Compaction
Abstract
Long-running AI agents often shorten their conversations by replacing older messages with a summary while keeping recent messages unchanged. When the model resumes, it normally processes these retained messages again, even though it has already computed and saved their internal key–value (KV) states. Reusing those states could save time, but the messages now follow a different history and appear at different positions. We study a simple reuse policy: process the new summary, move the cached keys to the retained messages’ new positions, and keep their cached values. To assess whether this shortcut preserves the agent’s next decision, we compare its replies with fresh computation of the same compressed input. We measure overall reply similarity, content agreement in both directions, and differences in proposed tool actions. In targeted Qwen3-8B replays, full-reply BGE-M3 similarity ranges from 0.780 to 1.000, yet replies that score as similar can propose different code edits. Separately, at measured native Mistral-Medium-3.5-128B compaction points, reuse makes the input-processing step 13.80× faster in geometric mean than the fastest measured fresh implementation, once the caches are prepared. Reuse therefore saves substantial computation in this measured setting, but leaves a semantic gap that task success or a single similarity score alone cannot characterize.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.