acceptodds
Under review as a conference paper at ICLR 2027

Watermark Basins: Black-Box Detection of Distilled LLMs

Abstract

Proprietary language models and their reasoning traces are vulnerable to API-based distillation, where an adversary trains a student model on the teacher model outputs. To provide evidence of distillation, model providers can watermark the teacher outputs, aiming to leave radioactivity in the student outputs. Existing black-box detectors search for such radioactivity using token-level statistics from the suspect model’s outputs. However, token-level signals are not always detectable; for instance, when the teacher model is distilled at a low temperature. In this work, we discover watermark basin, a persistent signal that enables more reliable watermark detection: during fine-tuning, the watermarked teacher traces embed a distinct, watermark key-specific geometric signal in the parameter space that lies orthogonal to the ordinary task-specific fine-tuning subspace. Building on this, we introduce Watermark Basin Alignment (WM-Basin), a black-box watermark detector that queries through the suspect model’s API, without requiring access to its weights or logits. In particular, WM-Basin maps the generated text into gradient representations using a fixed clean probe model, projects out the ordinary fine-tuning task directions, and verifies if the residual gradient is aligned with the watermark basin induced by the watermark key. We evaluate WM-Basin extensively across diverse teacher models and student architectures (spanning 135M to 8.2B parameters) distilled on mathematical reasoning. Under a strict 0.1% false-positive rate, WM-Basin detects 79% of distilled models at n-gram=1 and 45% at n-gram=3, substantially improving the state-of-the-art token-level baseline (29% and 1.3%, respectively). Furthermore, because the watermark signal resides in a persistent basin, it survives aggressive post-distillation modifications, including further fine-tuning, parameter quantization, pruning, and reinforcement-learning post-training

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.