OnDevMem: Efficient Long-Term Memory for On-Device LLM Agents
Abstract
Long-term memory is essential for LLM-based agents to preserve and reuse information accumulated over extended interactions. Agent memory has advanced substantially under cloud-based settings, whereas fully on-device long-term memory remains at an early stage. Deploying existing memory pipelines unchanged on resource-constrained devices can compromise memory quality and incur substantial construction costs from frequent LLM-based semantic processing. Memory maintenance can further increase foreground inference latency by competing for limited local resources. To address these challenges, this paper introduces OnDevMem, a cloud-independent memory framework that separates evidence preservation, semantic indexing, and maintenance scheduling. Topic-aware Evidence Memory (EM) preserves source interactions and provenance, while locally generated Semantic Signatures (SS) provide complementary semantic signals for retrieval without replacing the original evidence. Semantic enrichment at topic granularity reduces memory-construction cost by processing coherent interaction segments jointly. A resource-adaptive controller coordinates signature generation and index updates with foreground demand and device conditions to reduce foreground latency. Experiments and ablations across two local deployments on LoCoMo show that OnDevMem improves judge accuracy by 11.6%–14.5% over the strongest evaluated baseline and reduces memory-construction time by 88.4%–92.5% compared with the fastest baseline across deployments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.