Dataset Inference via Cumulative Evidence
Abstract
Dataset Inference (DI) aims to determine whether a given dataset was used to train a model. However, existing DI methods for large language models (LLMs) face two significant practical limitations. First, they assume gray-box access to token-level probabilities, whereas most commercial LLMs expose only generated text through black-box APIs. Second, they rely on rigid p-value hypothesis testing, which requires committing in advance to a specific dataset size and significance level. In practice, this makes auditing either too costly or statistically inconclusive, since adaptively adding additional samples invalidates standard p-value guarantees. We address these limitations with a fully black-box and sequential approach to DI. To overcome API restrictions, we introduce a novel method for estimating continuous per-token probabilities directly from discrete, label-only outputs, unlocking **DI in a fully black-box setting**. To address the rigidity of standard fixed-sample testing, we identify that e-values are particularly well suited for this setting. Based on e-values, we create a new sequential testing framework for DI. Crucially, when the initial dataset size is insufficient, our framework seamlessly accommodates additional data collection without violating statistical guarantees, enabling **DI with iterative evidence accumulation**. Together, these advances make DI practical for real-world LLM auditing. We validate our novel DI method on 6 diverse open-source LLMs with known training datasets and demonstrate its efficiency in production settings on 6 LLMs exposed via public APIs. Our cost analysis shows that, for instance, auditing GPT-5 via the OpenAI API using the WikiMIA-24 dataset with our method costs only $1.92.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.