acceptodds
Under review as a conference paper at ICLR 2027

SpanAudit: Benchmark Contamination Detection as Span Retrieval

Abstract

Benchmark contamination can cause evaluation scores to reflect prior exposure rather than the capabilities that a benchmark is intended to measure. Corpus-side detection can identify benchmark-specific information in training data before model training, but practical methods commonly rely on long exact matches, predefined retrieval units, or both. Rewritten benchmark content may no longer contain a sufficiently long contiguous -gram, while the relevant passage may occupy only part of a long document or cross a segmentation boundary. To address these coupled challenges, we introduce **SpanAudit** a query-conditioned method for approximate retrieval over variable-length source spans. **SpanAudit** keeps each complete benchmark item as a query, uses banded MinHash–LSH to identify spans with likely token-set overlap, and shares sketch computation across overlapping intervals instead of enumerating the quadratic span space. It then merges and localizes the retrieved candidates to produce inspectable evidence for downstream contamination adjudication. Across controlled experiments with progressively stronger paraphrases and increasingly long surrounding context, **SpanAudit** achieves the highest or tied-highest planted-pair recall in every reported setting, while retaining favorable latency among high-recall retrievers. Audits of three previously decontaminated instruction corpora also recover confirmed evidence absent from the other methods, including repeated answer and reasoning exposure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.