LAM: An Efficient And Controllable Lossy Agent Memory Framework
Abstract
Agent memory grows as agents read inputs, reason, and call tools. Longer histories increase inference cost and eventually exceed the context window. LLM-based summarization reduces this history but adds latency and provides no explicit bound on information loss. We propose LAM, a Lossy Agent Memory system with three components: a deterministic deduplication rule with a substitution bound on retrieval scores — a bound on score perturbation, not a certificate of unchanged ranking, a memory manager that preserves the cached prefix and overlaps compaction with inference, and a performance model that estimates compaction costs before deployment. On 600 agent trajectories, LAM removes 22.47% of observation tokens while retaining 99.984% of the measured gold-patch evidence. Evaluation on end-to-end latency shows that LAM achieves a 71.4×–91.6× end-to-end speedup from removing records before prefill, validating LAM's effectiveness and efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.