SpanAudit: Benchmark Contamination Detection as Span Retrieval
Abstract
Benchmark contamination can cause evaluation scores to reflect prior exposure rather than the capabilities that a benchmark is intended to measure. Corpus-side detection can identify benchmark-specific information in training data before model training, but practical methods commonly rely on long exact matches, predefined retrieval units, or both. Rewritten benchmark content may no longer contain a sufficiently long contiguous -gram, while the relevant passage may occupy only part of a long document or cross a segmentation boundary. To address these coupled challenges, we introduce **SpanAudit** a query-conditioned method for approximate retrieval over variable-length source spans. **SpanAudit** keeps each complete benchmark item as a query, uses banded MinHash–LSH to identify spans with likely token-set overlap, and shares sketch computation across overlapping intervals instead of enumerating the quadratic span space. It then merges and localizes the retrieved candidates to produce inspectable evidence for downstream contamination adjudication. Across controlled experiments with progressively stronger paraphrases and increasingly long surrounding context, **SpanAudit** achieves the highest or tied-highest planted-pair recall in every reported setting, while retaining favorable latency among high-recall retrievers. Audits of three previously decontaminated instruction corpora also recover confirmed evidence absent from the other methods, including repeated answer and reasoning exposure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.