acceptodds
Under review as a conference paper at ICLR 2027

Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory

Abstract

Memory plays a central role in enabling large language models (LLMs) to operate over sequential tasks by accumulating and reusing experience over time. However, existing evaluations of LLM memory mostly rely on aggregate metrics such as final hold-out accuracy or cumulative online performance. We argue that these metrics can be misleading: they collapse distinct memory behaviors into a single number and obscure critical failure modes such as forgetting and negative transfer. We introduce **SeqMem-Eval** to diagnose the behavior of sequentially evolving LLM memory. Drawing inspiration from continual learning, the framework targets a distinct test-time setting in which memory is external, prompt-mediated, and updated without changing model parameters. Beyond measuring whether the final memory state improves performance, **SeqMem-Eval** examines how memory states evolve, generalize, consolidate experience, and retain useful information during sequential inference. Specifically, it measures online utility, hold-out generalization, backward transfer, forgetting and efficiency, providing a finer-grained view of whether memory updates help current tasks, generalize to unseen tasks, improve past predictions, or degrade previously acquired knowledge. Through extensive experiments across diverse tasks and memory methods, we uncover several previously overlooked phenomena. In particular, we show that higher final or cumulative accuracy does not necessarily imply better memory quality: many methods exhibit strong performance gains while suffering from significant forgetting or negative transfers. Moreover, different memory designs exhibit distinct trade-offs between adaptability and stability, which are invisible under standard evaluation metrics. Our findings show that aggregate metrics systematically miss several recurring failure modes, suggesting that a multi-dimensional perspective is essential for understanding LLM memory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.