acceptodds
Under review as a conference paper at ICLR 2027

Gradient-Guided Radioactive Watermarking for Detecting Fine-Tuning on Teacher-Generated Data

Abstract

Model distillation transfers capabilities from powerful large language models (LLMs) to smaller downstream models, raising an important provenance question: Has a downstream model been fine-tuned on protected teacher-generated data? Existing methods embed watermark into generated text to detect such data use. However, their watermark statistics are evaluated on training-related or same-prompt contexts. In real-world scenarios, model owners may not know which data is used for distillation. To address this gap, we introduce Gradient-Guided Radioactive Watermarking (GGRW), a generation-time watermarking framework that links token selection to downstream fine-tuning dynamics. GGRW uses a proxy model to identify and sample tokens whose induced fine-tuning updates reinforce a private watermark direction, enabling detection on disjoint held-out contexts used in neither watermark construction nor student fine-tuning. Experimental results show that the watermark is detectable under dilution attacks and distillation, where student architectures are unknown. Furthermore, our method effectively detects the watermark on disjoint held-out contexts, whereas the baseline methods fail.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.