HiddenMark: Analytic Hidden-State Watermarking with Closed-Form Guarantees
Abstract
A provider deploying an LLM watermark has to commit to two numbers before the first token is generated: how much task accuracy the watermark will cost, and how much output must accumulate before detection is reliable. Existing watermarks answer both only after the fact, by running the watermark and measuring. We formalize the requirement as a Predictability Standard, give a certification procedure that applies to any watermark, and give a watermark designed to meet the full standard. HiddenMark perturbs generation along a single known direction in vocabulary space, derived from the LM head, so that its effect on every next-token distribution is available in closed form. This lets a deployer specify a budget that the controller enforces, and lets HiddenMark certify the outcome before generation. From one 30-prompt probe the procedure issues two certificates with finite-sample coverage guarantees, validated on HiddenMark, KGW and SynthID deployments: the output length needed to reach 95% detection (88% realized coverage at 90% nominal, within a factor of 2.0 over a 996-fold range) and a lower bound on the change in task accuracy. The detection certificate is as tight as the outcome can be measured: its error matches the disagreement between two independent 100-prompt evaluations of the same deployment. Under the single key a provider actually deploys, the usual closed-form threshold inflates the false positive rate by up to ; one precomputed scalar brings the inflation to –. Building the certificates exposed why watermarks destroy task accuracy while perplexity stays flat, and fixed the design: a spread-spectrum direction that bounds how far any token can be promoted. On GSM8K with a model that passes a qualification statistic computed from plain generation, HiddenMark matches KGW and SynthID on the accuracy–detection trade-off, and unlike them exposes a budget that is enforced during generation and a detector whose null holds under a deployed key; on a model the statistic rejects, it falls behind KGW.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.