A Memory You Can’t Recall Is No Memory at All: MEMENTO for Query-Agnostic KV Cache Eviction in Multi-Turn Inference
Abstract
Multi-turn agentic inference requires a bounded KV cache to support information needs that change across user requests and tool interactions. Repeatedly evicting entries according to the current query can specialize memory to the task at hand: once discarded, evidence needed by a later turn cannot be recovered by rescoring the surviving cache. We argue that useful memory requires both availability and accessibility: a memory the model cannot recall cannot support its response. We propose **MEMENTO**, a training-free streaming compression method that lets the model reread: it re-feeds each completed context chunk, and the attention of the reread supplies query-agnostic retention scores, using the content itself to generate retrieval cues. MEMENTO interleaves rereading with subsequent decoding in isolated batch rows of the same forward pass, amortizing scoring latency while maintaining a persistent cache across turns. A geometric analysis interprets reread queries as witnesses of retrieval cells and bounds local attention error on the regions they cover. Under explicit coverage and value-separation conditions, the same analysis yields a lower expected future-query error than a cache specialized to the current query at the same retention budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.