Before Answering: Evidence Sufficiency under Size-Matched Memory Construction
Abstract
Agents that answer questions from compressed or retrieved memory must recognize when the evidence a query needs is no longer in memory. Benchmarks for this task usually create insufficient-evidence examples by deleting supporting passages. We show that this construction leaks the label through memory size: on MuSiQue, a classifier that only counts paragraphs reaches an area under the ROC curve (AUROC) of for detecting unsafe memory, higher than the lexical estimator we initially evaluated. We propose a size-matched construction that provably removes this shortcut, and use it to study MemSafe, an estimator that cross-encodes the query with each memory unit and aggregates the units with a set transformer. Across three multi-hop question answering datasets and five seeds, MemSafe reaches and AUROC on MuSiQue and HotpotQA, to above a lexical baseline, while the third dataset, 2WikiMultiHopQA, is saturated. A frozen pretrained cross-encoder with a logistic head already closes of the MuSiQue gap between the lexical baseline and . At the same time, MemSafe degrades more than a weak baseline on the unanswerable questions released with MuSiQue, reaches only AUROC on SQuADÂ 2.0, and needs several thousand clinical training examples before it outperforms a feature-based estimator. Used as a gate for a 7B reader, it reduces the error rate on answered questions from to at coverage, outperforming both reader confidence and, on average, the ground-truth integrity label, although a 7B LLM judge is the better gate at coverage. These results indicate that the way insufficient evidence is constructed matters as much as the estimator that detects it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.