acceptodds
Under review as a conference paper at ICLR 2027

ReqMemBench: Benchmarking Coding Agents for Temporal Requirement-State Reconstruction in Evolving Projects

Abstract

Coding agents increasingly join ongoing software projects, where they must modify an evolving codebase in the context of a long-running client collaboration. Before acting, an agent must determine which requirements currently govern the task. This requires tracing how requirements have changed over time and determining which still apply, while distinguishing changes to the requirements themselves from evidence about their implementation. Existing coding and memory benchmarks offer limited visibility into whether agents correctly reconstruct the requirement state that should guide implementation. We introduce ReqMemBench, a benchmark for temporal requirement-state reconstruction built from real freelance project histories that have been rewritten for privacy and include a subset of role-played messages. At intermediate takeover points, ReqMemBench provides evidence-linked gold requirement states and evaluates four capabilities: selecting relevant historical requirements and evidence, reconstructing their pre-task states, updating requirements or identifying material clarification needs, and delivering the requested code changes. A two-phase protocol freezes history-based responses before agents choosing to act on eligible tasks receive repository access. Full History and Oracle Relevant History conditions assess the effect of history filtering and agents’ reconstruction accuracy when relevant evidence is explicitly supplied. Across seven configurations, agents over-select historical evidence and do not benefit from Oracle Relevant History; even the best RQ3 Post-State Score is only 0.1578, and no system passes one quarter of eligible RQ4 repositories. Reliable takeover therefore requires coherent temporal context and stronger update-or-clarify reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.