acceptodds
Under review as a conference paper at ICLR 2027

Top-k Probing for Black-Box LLM Watermark Detection

Abstract

Auditing an LLM watermark is hard when the model keeps its watermark key and token scores private. Existing black-box detectors usually examine responses under a fixed decoding setting; some also need carefully designed prompts that restrict the answer. We use a different clue: what happens when the model’s list of allowed tokens changes. Our top-kprobe sends the same prompt with a short and a long token list, then checks whether the model changes how it chooses between tokens on both lists. To the best of our knowledge, this is the first black-box LLM watermark test to use a controlled change in top-kon the same prompt as its probe. We build two detectors around this idea. Water-Zero uses an exact conditional test on ordinary text prefixes. It needs no answer template, trained detector, watermark key, or clean-model calibration. Water-Sonar learns patterns in the responses from each token list and covers more watermark families; it can use natural or crafted prompts. For natural prompts, we test held-out keys on C4 and OpenGen with three Qwen and Llama models. A separate crafted-prompt study covers nine mod- els and nine watermark families. Comparisons include plain-model audits, pooled exact tests, and tests that keep top-k fixed. Water-Zero detects watermarks that respond to a change in the allowed tokens, and its natural-prompt results transfer across the two datasets. Water-Sonar detects more watermark families, while its natural-prompt version remains effective for several major families. The fixed- top-ktests show that changing the token list creates the direct signal used by the exact detector.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.