Evaluating LLM Agent Behaviors and memory Under Evolving Constraints
Abstract
LLM agents are increasingly deployed as long-running assistants and companions equipped with memory systems. Yet, as real-world constraints inevitably change over time, no existing benchmark systematically evaluates whether an agent can maintain a correct picture of current rules across hundreds of sessions without falling back on stale information. We introduce RevokeBench, a benchmark of 1,202 tasks designed to evaluate agent compliance with dynamically evolving instructions over extensive histories. RevokeBench leverages a novel two-layer graph representation to track an evolving body of rules as unambiguous current states, enabling deterministic task annotation without relying on an LLM-as-a-judge framework. We construct 100 evaluation episodes across 6 realistic agent settings, each featuring massive context records of 300k–1M tokens and 10–15 graded tasks, drawn from an open pool of over 20,000 tasks. We evaluate 8 backbone models, further pairing 2 selected models with 4 memory frameworks. Our results show that state violation rates range from 31% to 83% across flash-tier models, and frontier models do not uniformly outperform their flash-tier counterparts. Across all four memory frameworks evaluated spanning middleware and self-evolving memory, agents incur more violations and complete fewer tasks than a simple heuristic baseline. While our introduced Rule-Tracking Memory baseline improves performance, reliable constraint tracking in long-horizon deployments remains an open challenge. We release the full dataset, task generator, solver, evaluation framework, and baselines to support future research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.