acceptodds
Under review as a conference paper at ICLR 2027

Does Training on a Text Make It More Likely? Not Under Distillation

Abstract

Likelihood-based data audits assume that a model assigns higher probability to text it was trained on. We test this assumption with a randomized-inclusion design, changing the prediction target while keeping the input prefixes fixed, and find that it fails under distillation. We randomly choose which problems enter training, score every trajectory before and after training, and compare the gain of the included trajectories with the gain of the excluded ones. Fine-tuning on the original tokens gives the included trajectories the larger gain. Using the teacher's argmax labels instead, however, reverses this advantage: the contrast in gains changes from +0.016 to −0.019 nats per token, and the reversal replicates in a second model family. It arises at the roughly 5% of positions where the teacher's label differs from the original token: training only at these positions reproduces the negative contrast, whereas training only at the remaining positions yields a positive one. To first order, the expected contrast is non-negative when training on the scored tokens and can take either sign when the target changes. For auditing, this reversal flips a common membership score, which ranks texts by how much their likelihood rose: the score separates included from excluded trajectories after fine-tuning and ranks them the wrong way round after distillation. A control in which the student is distilled from a copy of itself reverses the score just as much despite a far smaller contrast, so a reversed score alone cannot identify distillation. A randomized contrast in either direction shows that data was used; a null does not show that it was not.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.