When the Probe Is the Problem: Black-Box Watermark Detection on a Deployed LLM
Abstract
Black-box watermark presence tests ask whether a language model carries a watermark using samples alone. We audit the Red-Green test on Claude, whose provider announced a SynthID-Text deployment without disclosing its parameters. Across ten initial configurations the test never rejects, despite detecting a local reference at comparable sample sizes. Seeding the reference's repeated-context history with the prompt blinds the echo probe while leaving the watermark active elsewhere. Two redesigns expose native word preferences and response-planning confounds, demonstrating the need for controls from the deployed model's family. A long assembled-context probe, scored within each prefix, produces a substantial shift on Fable absent on two presumably unwatermarked siblings. Its primary statistic rejects, but the pre-registered decision is inconclusive because refusals are concentrated by context. Follow-ups do not establish a decisive replication, and the shift's watermark origin remains unresolved. We provide an auditing checklist and release code, prompts and parsed responses. The results show how a deployment can violate a presence test's assumptions, and why statistical rejection and a null each require interpretation against explicit controls.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.