acceptodds
Under review as a conference paper at ICLR 2027

MemEvoBench: Feedback-Guided Case-Type Evolution for Long-Term Memory Evaluation

Abstract

Existing long-term memory benchmarks typically organize evaluation around predefined case types, making it difficult to incorporate requirements exposed by observed system failures. We introduce MemEvoBench, a feedback-guided benchmark evolution framework grounded in de-identified multi-session user dialogues. Each case type specifies a target memory behavior and executable rules for evidence discovery, case generation, and verification. The framework uses validated cross-system evaluation feedback to evolve these specifications and incorporates admitted types into subsequent case mining, turning instance-level failures into reusable evaluation requirements. In a ten-round study over six user histories, MemEvoBench constructed 1,553 evidence-grounded cases and expanded four seed types into 11 consolidated types. The evolved types exposed failures both in systems used during evolution and in additional systems that supplied no feedback. These findings support feedback-guided case-type evolution as a complement to predefined memory benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.