CMB: A Coding Memory Benchmark for Long-Horizon and Interleaved Coding Agents
Abstract
Coding agents increasingly operate in long sessions where tasks are interrupted, requirements evolve, branches diverge, and unfinished work must later be resumed. Memory must therefore support subsequent repository actions, not merely recall prior text. We introduce CodeMemoryBench (CMB), an execution-grounded benchmark for memory in long-session coding agents. CMB constructs controlled, interleaved task histories in real repositories, freezes interaction streams and repository checkpoints, and replays identical histories to every memory mechanism. It comprises 216 code-editing tasks and 750 diagnostic questions over 5,595 turns from nine repositories in four languages. Experiments with five existing memory mechanisms reveal substantial room for improvement: the strongest mechanism resolves only 36.6% of CMB tasks with the stronger of two agent backbones. To address this limitation, we propose a three-layer memory that links task facts, reusable experience, and concrete code diffs through provenance and evolution relations. It reaches 43.1% and 35.6% on CMB across the two backbones, compared with 36.6% and 28.2% for the strongest baselines. On SWE-ContextBench, it achieves 47.5% and 28.3%, versus 45.5% and 25.3% for the strongest competing memories. CMB connects diagnostic recall with executable outcomes, measuring whether memory enables agents to resume and complete coding work.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.