SOLM: SELF-ORGANIZING LATTICE MEMORY FOR LLM AGENTS
Abstract
Long-term memory for LLM agents requires imposing structure on an evergrowing history, but several recent memory architectures pay for that structure with real, per-message LLM calls: A-Mem generates a structured note and re-evaluates evolution for a set of neighboring memories on every insertion, MemTree re-aggregates content across an average of over three ancestor nodes per insertion, and Zep extracts entities, relations, and deduplicates edges via multiple LLM calls per episode. This construction cost scales with memory volume, tying a deployment’s throughput directly to model inference capacity. SOLM’s structure-building (bottom-up placement, quantization-error-driven splitting, cross-branch weak links) needs zero LLM calls, computed from embeddings alone. Its only LLM usage: one fixed call per query for generation, and one optional, cheap small-model call per message for importance scoring; this is the sole cost scaling with volume, far lighter than A-Mem, MemTree, or Zep’s perinsertion costs. Weak links directly address what we call aggregation dilution: a cluster’s centroid can fail to resemble a query even when a member buried inside it would match well, causing that whole branch to be wrongly discarded. (A nongenerative reranking step narrows and reorders an already-retrieved pool at query time and is not counted here; see Section 3.5.) On LongMemEval-S, a controlled ablation (mean over 3 independent runs) shows weak links yield a relative accuracy improvement of 7.7% (68.6% to 74.0%) at matched settings; at nearly identical retrieved-context size, SOLM outperforms a flat similarity-retrieval baseline by a relative 13.8% (mean over 3 runs: 76.0% vs. 66.7%); and accuracy scales predictably with the retrieved-context budget: 76.0%, 83.0%, and 86.0% when the top-5, top-10, and top-15 retrieved memories, respectively, are used as answer context. These results show that competitive long-term memory accuracy does not require scaling LLM usage in proportion to memory volume.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.