MemEvoBench: Feedback-Guided Case-Type Evolution for Long-Term Memory Evaluation
Abstract
Existing long-term memory benchmarks typically organize evaluation around predefined case types, making it difficult to incorporate requirements exposed by observed system failures. We introduce MemEvoBench, a feedback-guided benchmark evolution framework grounded in de-identified multi-session user dialogues. Each case type specifies a target memory behavior and executable rules for evidence discovery, case generation, and verification. The framework uses validated cross-system evaluation feedback to evolve these specifications and incorporates admitted types into subsequent case mining, turning instance-level failures into reusable evaluation requirements. In a ten-round study over six user histories, MemEvoBench constructed 1,553 evidence-grounded cases and expanded four seed types into 11 consolidated types. The evolved types exposed failures both in systems used during evolution and in additional systems that supplied no feedback. These findings support feedback-guided case-type evolution as a complement to predefined memory benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.