acceptodds
Under review as a conference paper at ICLR 2027

MemGym: a Long-Horizon Memory Environment for LLM Agents

Abstract

Long-horizon LLM agents must keep deciding which observations and tool outputs to hold in context. Most memory benchmarks test recall in multi-turn chat. Agent benchmarks where these decisions arise report only end-task success, which mixes memory failures with reasoning and tool-use failures. We present MemGym, an agentic memory benchmark that inserts a memory module between the interaction history and the reasoning model. Its six tracks range from tool-use dialogue and deep-research search to coding and computer use. Each track supports controlled comparisons of memory modules against a track-specific no-memory reference under a fixed environment, scaffold, and reasoner. Four tracks wrap agent benchmarks where forgotten information is often recoverable from the environment. We therefore construct MemGym-CodeQA and MemGym-DR at controlled context lengths with construction-time checks of fact recoverability and answer support. To make SWE-Gym comparisons cheaper, we train MemRM, a 1.7B-parameter scorer that predicts whether the next action after a compression returns the recorded reference output. This local check costs one reasoner step and one classifier call. Full rollouts still measure end-task outcome. We evaluate eight memory strategies with four reasoning models. A rolling summary does not improve SWE-Gym performance. Both summary strategies yield higher mean scores on tool-use dialogue, and a structured summary does so on web tasks. ProgramBench episodes often outgrow the context window. A structured summary raises their mean fraction of hidden tests passed from 9% to 49% under the same step budget. We release MemGym and MemRM with the labeled trajectory corpus and pipeline code at https://anonymous.4open.science/r/MemGym-003C.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.