MemAgent-Critic: Adaptive Self-Supervised Critic Training for Long-Context Memory Agents
Abstract
Training long-context memory agents with reinforcement learning faces sparse feedback: a single verifiable reward arrives only after the entire document is processed. A learned memory-trace critic supplies additional trajectory-level guidance, but a fixed critic can become misaligned with the evolving policy. To adapt this guidance to current-policy trajectories, we propose **MemAgent-Critic**, which updates a 1.5B critic online alongside a 7B memory policy. After initialization with GPT-5.4- and policy-generated traces, the critic is updated using within-question comparisons of combined answer and critic rewards. These updates reuse existing policy rollouts without additional teacher annotations. On RULER-HotpotQA, the resulting 7B policy achieves **93.75%** accuracy at 7K tokens and **80.47%** at 3.5M tokens. Across ten context lengths, the adaptive variant obtains an average accuracy of **88.75%**, compared with 87.27% for the frozen-critic variant, although the variants use different training durations and the frozen critic performs better at the two longest lengths. Additional evaluations report gains on several document QA, summarization, and code-completion tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.