acceptodds
Under review as a conference paper at ICLR 2027

Does Distillation Leave a Trace? Source Detectability in Complex Post-Training Pipelines

Abstract

Model post-training pipelines routinely distill data from multiple source models, across different domains and over multiple stages of training. Without proper reporting, it becomes difficult to determine which sources contributed to a released model, raising concerns around IP enforcement and safeguard circumvention. This raises a fundamental question: _do complex post-training decisions affect distillation source detectability_. We present the first systematic study of this question. We build a controlled testbed of 169 models distilled from ten source models across 14 settings, varied across different post-training design decisions. For each source, we test whether the model distilled from it shows greater similarity to that source than reference models trained on alternative sources. We find that while detectability magnitude varies by source and domain, it is largely preserved across post-training choices such as multi-domain distillation, target-model changes, and further on-policy distillation. One key exception: mixing multiple sources within the same domain generally degrades detectability for involved sources, though cross-domain mixing does not. Finally, under more realistic audit assumptions, where reference models do not exactly match the distilled model's pipeline, we find that signals based on the distilled model's likelihoods prove fragile, while stylistic signals remain robust. We release our models, data, and code to support future work on distillation detection and improving our understanding of post-training mechanisms.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.