Agent Memory That Pays Its Own Rent
Abstract
An agent can keep each memory item as text or as its key-value (KV) cache. Text is prefilled again at every read; the cache skips that prefill but pays for storage while it is kept. Serving engines such as SGLang and Mooncake keep the cache, and agent memory systems such as MemGPT and Mem0 keep text. The cheaper choice depends on the deployment. At the same hourly price, the reuse needed for caching to pay off differs 26.3 times between two accelerators. We present METER, a controller that derives a break-even period, over which storing a cache costs as much as recomputing it, from model configuration, storage prices, and measured prefill and load times. METER admits an item once two of its reads fall within one period; each read then renews the cache for another period, after which it lapses to text. On LongMemEval and LoCoMo, across 240 combinations of reader model and memory lifetime, METER costs at most 1.51 times an offline optimum that knows all future reads, and caches exactly the items that optimum caches. Never caching reaches 1.94 and METER without admission 1.99; rules that ignore storage prices reach 160 to 3,191. Twenty-one executed runs on two accelerators reproduce the item selection. On Mooncake serving traces, with over a thousand times more reuse, METER without admission costs 1.01 against 2.23 for never caching, and is provably within twice the cheapest choice on every gap between reads.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.