acceptodds
Under review as a conference paper at ICLR 2027

WMMD-Bench: Do LLM Ownership Watermarks Survive Model Distillation?

Abstract

Modern LLM watermarking and fingerprinting methods are increasingly used to protect model ownership, yet their robustness to model distillation remains insufficiently understood. In a model extraction setting, an adversary can distill a protected teacher into a fresh student using only teacher supervision, potentially weakening or removing the ownership evidence embedded in the original model. We introduce WMMD-Bench, a systematic evaluation of ten recent active and passive ownership mechanisms under direct hard-label distillation, semantic-preserving response transformation, online soft-logit distillation, and same- and cross-lineage student models. Across native detection metrics, clean controls, utility evaluation, applicability audits, and five 11-checkpoint training trajectories, we find that the evaluated active watermarks are not reliably preserved under model distillation. Their registered signals are substantially weakened or remain absent in distilled students, while richer soft-logit supervision provides only limited and method-dependent recovery. Training trajectories further suggest that this vulnerability is primarily associated with limited watermark acquisition, rather than strong acquisition followed by later forgetting. Passive fingerprints appear substantially more stable, but all five already produce positive detections on clean same-lineage students before distillation, indicating that their apparent robustness can be confounded by pre-existing model similarity. Overall, our results reveal model distillation as a substantial weakness of current active LLM watermarking methods and highlight the need for ownership mechanisms explicitly designed to survive knowledge transfer into independently trained student models. Code and evaluation artifacts are available at [https://anonymous.4open.science/r/WMMD-Bench-FCA6/](https://anonymous.4open.science/r/WMMD-Bench-FCA6/).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.