acceptodds
Under review as a conference paper at ICLR 2027

Two Counts over Text: What a Memory Benchmark Can See, Measured Before Anything Is Run

Abstract

A long-term memory benchmark is cited for an architectural distinction: a system that maintains a composable state against one that indexes what it was told. That citation is never checked against the benchmark's own text, although it can be. Where an item's gold answer is announced verbatim inside a single logged record, returning that one record suffices and no composition is being tested — but auditing this today means running systems, a judge and an API bill, so it happens after a benchmark ships rather than before. We build an instrument that needs no model, no judge and no API call: two counts over text and a retriever. V is the share of items whose gold answer appears verbatim in one record; R-str(k) is the share of that containment gap a budget-k retriever fails to close, and k* is the budget where it closes, the audit's one free constant swept and priced within 0.008 of optimal. A proved proposition makes the instrument one-sided: it can condemn a task as blind and can never acquit one. Run blind on four corpora, k* recovers class labels fixed by role before any dataset was named — single-hop dimensions close at 8, 8 and 16, composition dimensions do not close at 32, and a needle task returns the k* of 1 a working retriever must return. On a flagship benchmark built to be immune to context stuffing, V runs 0.55 to 0.64 on knowledge_update against 0.00 to 0.07 elsewhere: more than half its checkable update answers are announced verbatim in one turn, and across a 74x range of that benchmark's own headline axis the counts barely move. Priced on a corpus built to defeat it, V travels while k* does not, since k* also measures a reader's lexical reach. The reporting rule is one-sided too: a high V condemns, a low reader score proves nothing, and authors can publish V and k* per dimension before shipping, at the budget the baseline was allowed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.