Batch size is not sample size: near-tie amplification in contrastive membership inference
Abstract
Dataset inference asks whether a set of examples was used to train a model. At web scale each example carries almost no evidence, so audits average a per-example score over the set and rely on a square-root law: separability grows as the square root of the number of examples. We show that this law fails for contrastive image–text models when the score is the InfoNCE loss computed within a batch, because batch size is then not sample size: a larger batch averages more pairs but also gives every pair more negatives to beat. On a CLIP model trained on 400M pairs with a pre-training holdout, the effect size of the batch-mean loss at the trained temperature grows as , whereas scoring the same pairs against fixed negatives restores . The mechanism is near-tie amplification: the loss mean shift increases when a true pair barely beats its hardest negative; the gain over therefore peaks at an intermediate number of negatives and then reverses. Fixed-negative measurements alone predict the in-batch curve up to batch size on three separately trained checkpoints and, in a pre-registered test, the optimal number of negatives and the detector ranking on a 40M-pair checkpoint of the same recipe. Because the signal lives in a pair's own similarity, auditors need no negatives beyond their reference set and no temperature: the expected crossing score obeys the square-root law and, when membership mainly affects poorly aligned pairs, outperforms the in-batch loss by up to 11 AUC points on 1,024 pairs and raises the true-positive rate at 1% false-positive rate by up to 38 percentile points (57% vs. 95%). Audits remain fragile: verdicts need thousands of pairs, can become uninterpretable when fewer than about a fifth of candidates are members, and can be manufactured by deleting of the least-aligned non-member pairs. We recommend blind-model controls and pre-registered candidate sets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.