Gradient-Guided Radioactive Watermarking for Detecting Fine-Tuning on Teacher-Generated Data
Abstract
Model distillation transfers capabilities from powerful large language models (LLMs) to smaller downstream models, raising an important provenance question: Has a downstream model been fine-tuned on protected teacher-generated data? Existing methods embed watermark into generated text to detect such data use. However, their watermark statistics are evaluated on training-related or same-prompt contexts. In real-world scenarios, model owners may not know which data is used for distillation. To address this gap, we introduce Gradient-Guided Radioactive Watermarking (GGRW), a generation-time watermarking framework that links token selection to downstream fine-tuning dynamics. GGRW uses a proxy model to identify and sample tokens whose induced fine-tuning updates reinforce a private watermark direction, enabling detection on disjoint held-out contexts used in neither watermark construction nor student fine-tuning. Experimental results show that the watermark is detectable under dilution attacks and distillation, where student architectures are unknown. Furthermore, our method effectively detects the watermark on disjoint held-out contexts, whereas the baseline methods fail.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.