acceptodds
Under review as a conference paper at ICLR 2027

MemGround: Evaluating Memory Behavior of Large Language Models in Long-Horizon Gamified Scenarios

Abstract

Long-term memory in large language models (LLMs) is typically evaluated by asking whether a model can retrieve and reason over information from a fixed interaction history. However, real interactive agents face a fundamentally different challenge: relevant information is progressively revealed, world states continuously evolve, and useful evidence must be discovered, organized, updated, and reused over extended interactions. This suggests evaluating long-term memory not merely as information retention, but as memory behavior—the ability to organize and apply accumulated evidence throughout interaction. We introduce MemGround, a benchmark that operationalizes this perspective through long-horizon gamified interactive environments. MemGround assesses memory behavior across three memory-structure patterns: Linear Memory, Temporal Memory, and Graph-Structured Memory. To evaluate memory behavior comprehensively, we design a two-stage framework that combines end-to-end behavioral evaluation and process-level diagnosis. The end-to-end evaluation examines whether models can use accumulated information to answer questions correctly, make sustained progress during interaction, recover temporal organization among fragmented memories, and navigate structural dependencies coherently. We quantify these behaviors using Question-Answer Overall (QA Overall), Memory Fragments Unlocked (MFU), Memory Fragments with Correct Order (MFCO), and Exploration Trajectory Diagram (ETD). Beyond these observable outcomes, we further introduce Memory Score (MS) and Reasoning Score (RS) to identify whether failures stem from insufficient evidence retrieval or from incorrect reasoning over retrieved evidence. Experiments across state-of-the-art LLMs and memory-augmented agents reveal persistent limitations in dynamic state tracking, temporal association, and evidence-grounded reasoning over long-horizon interactions. We hope MemGround can serve as a systematic testbed for advancing the memory capabilities of future LLMs and agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.