acceptodds
Under review as a conference paper at ICLR 2027

Before Answering: Evidence Sufficiency under Size-Matched Memory Construction

Abstract

Agents that answer questions from compressed or retrieved memory must recognize when the evidence a query needs is no longer in memory. Benchmarks for this task usually create insufficient-evidence examples by deleting supporting passages. We show that this construction leaks the label through memory size: on MuSiQue, a classifier that only counts paragraphs reaches an area under the ROC curve (AUROC) of for detecting unsafe memory, higher than the lexical estimator we initially evaluated. We propose a size-matched construction that provably removes this shortcut, and use it to study MemSafe, an estimator that cross-encodes the query with each memory unit and aggregates the units with a set transformer. Across three multi-hop question answering datasets and five seeds, MemSafe reaches and AUROC on MuSiQue and HotpotQA, to above a lexical baseline, while the third dataset, 2WikiMultiHopQA, is saturated. A frozen pretrained cross-encoder with a logistic head already closes of the MuSiQue gap between the lexical baseline and . At the same time, MemSafe degrades more than a weak baseline on the unanswerable questions released with MuSiQue, reaches only AUROC on SQuAD 2.0, and needs several thousand clinical training examples before it outperforms a feature-based estimator. Used as a gate for a 7B reader, it reduces the error rate on answered questions from to at coverage, outperforming both reader confidence and, on average, the ground-truth integrity label, although a 7B LLM judge is the better gate at coverage. These results indicate that the way insufficient evidence is constructed matters as much as the estimator that detects it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.