Executable Gold Semantics: Compiling Gold Labels for Auditing Memory Transitions in Language Agents
Abstract
Memory systems let language agents carry information across long interactions, and they are evaluated by whether the agent eventually answers questions correctly. Whether each individual memory operation was right is rarely measured, because per-operation labels are held to be impractical to obtain. The one operation-level benchmark generates its memory points with a language model and verifies them by hand, and no benchmark records how a displaced belief should be read, who observed a change, or what the agent still owes. We show that once a scenario schema is fixed, these labels can be compiled instead of annotated. A small router of belief-change operators computes, for every event, the correct operation, the successor memory, whether the displaced belief was once true or wrong all along, who observed the change, and the open obligations. We prove that the router is well defined and runs in polynomial time, that every compiled trace satisfies contracts stated over states alone, and that six axioms characterize it, so that dropping any one yields an operator with a recognizable design flaw. The same semantics generates gold with paired counterfactuals, judges systems against it, and defines a training signal and a runtime guard. We release MEMLEDGER, the compiler and a benchmark of 900 sessions in two languages, and use it to ask what each evaluation view can see. On a balanced suite of 1,200 seeded contract violations, questions about final values detect under one percent, the best single question per violation class detects three quarters, and one large language model given the complete record in a single pass detects under ten percent. Changing only how a displaced belief is interpreted inside a real memory framework moves its answers about the past from 2 to 33 correct out of 36. Two author annotators, working without access to the router or to each other's labels, agreed with the compiled labels on 1,199 and 1,177 of 1,200 items, and with each other on 1,178.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.