acceptodds
Under review as a conference paper at ICLR 2027

FlashMEM: A Lightweight Memory System for Language Agents

Abstract

Long-running language agents repeatedly draw on growing interaction histories. Memory construction organizes history into compact context, but summaries formed before future queries are known may omit needed details. More elaborate construction and maintenance also incur additional model calls, repeated encoding, and persistent state storage. We revisit the role of constructed memory: it should provide directly readable, organized context and help locate relevant sources, rather than serve as the sole carrier of historical evidence. We introduce , which co-designs agent-side memory representation and server-side key–value (KV) cache reuse. On the agent side, one-pass construction produces compact, source-linked memory. Answering combines retrieved memory with a small amount of original evidence from its linked sources. On the server side, reuses recurring context across construction calls and retains KV states for compact memory rather than the full history across queries. Across three datasets, \ours achieves relative F1 gains of 16.1%, 33.1%, and 41.5%, while reducing construction costs by 97.0% and inference costs by 92.2%. \ours's server-side optimizations reduce time to first token (TTFT) by 66.8%, selection overhead for KV recomputation by 99.3%, and long-term PIC KV storage costs by 84.9–93.8%. The framework also accelerates other agent memory methods. We will release our code after anonymous review.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.