MemReq: Contract-Conditional Requirement Profiles for Diagnosing Agent Memory Failures
Abstract
End-to-end memory scores do not identify whether a named operation was exercised, its prerequisite state was retained, or evidence reached the Reader. We introduce MemReq, an offline diagnostic protocol that freezes the workload, information, Reader, resource, and scoring contract; contrasts executed operations; probes prerequisites; audits frozen-state evidence; and tests a bounded, gold-free intervention. Its requirement profiles are empirical and contract-conditional, not workload-intrinsic necessities. On 58 shared multi-hop questions with byte-identical contexts and matched 16,384-token completion caps, Qwen35 and GPT-OSS-120B yield V×C interactions of +27.59 and −5.17 percentage points, respectively. Reader identity and reasoning settings both differ, so this is not a model-only causal effect. In a crossed Mem0-RV diagnostic, 62/64 instances with a valid update opportunity reach latest-only state: apparent version failure can originate before replacement is possible. On frozen stores, version filtering improves Additive and A-MEM under dense and graph readout but leaves RV near zero; RV’s tested graph contexts do not change under filtering. The graph intervention does not reproduce its 6k gain on the official 32k-history configuration, where query anchoring limits delivery. MemReq therefore localizes observable failures without inferring unexercised primitives or universal architecture rankings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.