Not Yet a Measurement: Instrument Dependence in LLM-Judge Audits of Agent Memory
Abstract
Lifelong-memory agents write durable records about their users that later sessions read as background fact, without access to the transcript that produced them. Prior work has studied the read side; the write objective, and the provenance of what it admits, has not been a controlled variable. We set out to measure how often such entries assert specifics the user never stated, and report instead that this is not yet a measurement one can trust. On a controlled testbed over LongMemEval we vary the write objective, the elicitation arm and the agent model, then audit the same entries with two LLM judges under two defensible instruments, two administration modes, and criteria-based coding with rules committed before any entry was labelled. The answer depends on the instrument more than on the data. On identical entries, estimated prevalence moves from 2.6% to 41.0% at chance agreement (κ = 0.011). The two instruments rank the write objectives differently, yield different multiplicity-surviving contrasts, and reverse the sign of the elicitation effect, while agreeing at κ ≈ 0.48 within the same judge. The disagreement is directional rather than noise: given the same rules and worked examples in-prompt, judges resolve borderline entries toward "the user said this," the direction that makes a pipeline look cleaner than criteria-based coding says it is. Two things did hold still, and both are contrasts rather than rates. Administration mode — one entry per call versus a shuffled batch, with the evidence and the question unchanged — moves inter-judge κ by +0.209 [0.139, 0.281]; measured again with a different judge pair the difference is +0.213, while the absolute agreement levels move by 0.10 between pairs. We have not seen this design choice reported. The second is the model doing the writing. These results are negative about a measurement, not about the phenomenon. No human validated any label, so our κ values index instrument stability rather than accuracy and we report no prevalence; a human-coded calibration set is the prerequisite we identify. We close with a reporting checklist, each item earned by one of our results.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.