acceptodds
Under review as a conference paper at ICLR 2027

Uncovering the Tells: Efficient Black-Box Audits for LLM Data Contamination

Abstract

Existing methods for assessing training data contamination in Large Language Models achieve high accuracy, but fundamentally rely on gray-box access, rendering them inapplicable to fully black-box environments. This limitation forces auditors to rely on full-sequence sampling or rigid heuristics, both of which incur prohibitive inference costs and struggle to isolate the underlying memorization signal. To overcome this, we propose a generation-based black-box framework that infers context-induced information gain from targeted token discrepancies. We formalize a budget-constrained auditing regime, building on the empirical observation that the membership signal is highly sparse and concentrated in a few highly informative tokens. We statistically formalize this sparsity, demonstrating that targeting these anomalies avoids the severe sample complexity penalty of uniform auditing, allowing for accurate detection without probing the full document. Empirically, our approach achieves state-of-the-art black-box separation accuracy while needing about fewer queries than uniform position sampling. Finally, we apply this methodology to closed frontier models, exploring its capacity and fundamental limits in surfacing potential training data overlaps on commercial APIs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.