Generation–Evaluation Coupling in LoCoMo: A Construct-Validity Audit of Long-Term Memory Evaluation
Abstract
Long-term-memory benchmark scores can depend on both answer-generation instructions and how evaluators recognize the resulting behavior. We audit this interaction in LoCoMo with one fixed answer reader; within each of four retrieval/memory backends, questions, retrieved evidence, and decoding are held fixed across prompt conditions. Changing the generation prompt increases our archived all-item score by 0.214 on average. On unanswerable questions, a targeted change in the abstention instruction increases lexical credit by 66.7 percentage points while blinded human-judged abstention increases by 25.9 points, leaving a 40.8-point measurement gap (95% CI: 34.7–46.8). In the broader comparison between a neutral and a benchmark-aligned prompt on answerable questions, archived token F1 increases while GPT-4o reference-answer correctness decreases by 5.26 percentage points; a stratified human study on 240 paired units shows the same direction on the matched sample. A semantic abstention judge reduces lexical under-recognition but introduces false positives. These findings show that generation instructions can change both response behavior and its measurement. We recommend specifying generation and scoring protocols jointly and reporting answer correctness and abstention separately before interpreting aggregate gains as improvements in memory capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.