The Signal Beneath the Words: Paraphrase-Resistant Watermarking for LLM Distillation
Abstract
Unauthorized distillation poses a growing threat to the intellectual property of large language models. Existing tracing methods commonly embed watermarks through token-level statistical biases that can be inherited by student models. Yet even simple paraphrasing before distillation can substantially weaken this inheritance: paraphrasing readily rewrites lexical choices while preserving the useful content of the response. This exposes a fundamental weakness of token-level watermarks: their carrier lies precisely in the part of language that paraphrasing is designed to change. We take a different perspective: rather than watermarking what words a model chooses, we watermark a faint but persistent affective tendency in how it speaks. Our key observation is that language carries more than explicit content. Responses can exhibit a subtle emotional undertone that accompanies the text without dominating its meaning or compromising response quality, much like a weak aspect of personality. When expressed consistently across teacher responses, such an affective tendency can become part of what a distilled student learns. Based on this observation, we propose an emotion-circuit watermarking framework that implants a weak, quality-preserving affective tendency, e.g., a slight tendency toward disgust, by steering emotion-related internal activations during teacher generation. The watermark remains subtle in individual responses but becomes statistically detectable across outputs. Verification determines whether this tendency has been inherited by comparing the student's target-emotion responses and their contrast with other emotions against matched unwatermarked reference students. Across two student families and five token-watermark baselines, our watermark remains detectable in 100% of students after neutral paraphrasing of the teacher data, including when distillation questions come from a second source.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.