acceptodds
Under review as a conference paper at ICLR 2027

Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks

Abstract

LLM-based assistants increasingly rely on memory extracted from past conversations to serve users over time. However, real-world assistants face diverse tasks requiring different types of memory, which existing work largely ignores. We formalize the heterogeneous memory extraction task and introduce BEHEMOTH, a benchmark that repurposes 18 existing datasets spanning personalization, problem-solving, and agentic tasks, using a downstream utility-driven metric for systematic evaluation. Our empirical analysis confirms that no single static extraction prompt is optimal for all task categories, and that existing self-evolving prompt optimization frameworks degrade when training tasks are heterogeneous. To address this, we propose CluE, a self-evolving method that restructures analysis hierarchically: training examples are partitioned into clusters, each cluster is analyzed independently to keep feedback concentrated, and cross-cluster insights are synthesized into a single prompt update. Experiments on BEHEMOTH show that CluE consistently outperforms prior self-evolving frameworks. Together, the task formulation, benchmark, and method provide a foundation for memory extraction across heterogeneous real-world tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.