acceptodds
Under review as a conference paper at ICLR 2027

DistillTrace: Tracing LLM Distillation via Keyed Solution Preferences

Abstract

Large language model outputs can be collected for unauthorized distillation, motivating providers to determine whether a student was trained on their answers. Yet a watermark detectable in teacher outputs may not survive distillation. Practical tracing therefore requires a learnable signal, broad task coverage, and preserved teacher utility. We introduce DistillTrace, which encodes consistent solution preferences as a watermark. Two LoRA adapters learn response families from quality-filtered, self-distilled answer pairs. Keyed semantic hashing assigns one family to each predicted question category, giving related questions a consistent preference that students can learn through ordinary distillation. A fixed private detector supports black-box tracing and selects watermarked answers from other generators without retraining. Construction and embedding require neither proxy students nor token-level decoding biases. In a matched Qwen-to-Llama comparison on MBPP, aggregating 256 responses yields 99.88% AUROC and 97.77% TPR against matched clean students. Teacher utility is maintained or improved on all five tasks evaluated in this comparison. Separately, question-count-weighted teacher utility rises from 53.73% to 54.50% across four benchmarks. In multi-domain distillation, the semantic-hashing configuration achieves stronger aggregate student detection than the non-semantic configuration. Across four generators and three tasks held out during adapter and detector training, online selection from four candidates achieves 95.52%-100.00% aggregate watermark-detection TPR over 256 responses, extending embedding without additional training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.