acceptodds
Under review as a conference paper at ICLR 2027

Beyond Surface Tokens: Watermarking Language Models through Private Solution Choices

Abstract

Watermarking provides evidence of model ownership, but signals carried by wording and code patterns are exposed to rewriting and normalization. We introduce, to our knowledge, the first watermark carried by which valid solution a model returns to a constraint problem, rather than by how the answer is worded. A private key selects one of two verified solutions, and joint fine-tuning combines these protected targets with ordinary problem-solving examples. Verification compares parsed solutions through the model's input-output interface, so rewriting that preserves the solution preserves the mark. Removing the mark while keeping answers correct requires another valid solution, which links removal to the Another Solution Problem (ASP), NP-complete in the worst case for the domains we use. Across X3C, Kakuro, and graph coloring, joint training raises accuracy on 1,536 unseen problems from 16.21% to 42.45%, while producing the exact owner solution on all 96 protected problems used in training. General benchmark performance remains close to the base model. Given WM's valid answers, five commercial reasoning models find different valid solutions on 0% to 20.83% of the problems under the reported inference budgets. Ordinary continued fine-tuning retains nearly all owner matches, whereas targeted interventions reduce owner matches only together with unseen-problem accuracy, a cost that general benchmark scores alone do not reveal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.