Constitution-Guided Watermarking
Abstract
Watermarking enables language model providers to identify text generated by their models. However, its desired properties can conflict (i.e., stronger watermark signals can degrade text quality), while designs that resist editing may also facilitate forgery. Providers address these trade-offs by choosing configurations that balance competing objectives or prioritize particular properties. Either approach imposes a shared operating point on requests with different requirements, potentially sacrificing quality where wording preservation matters or robustness where reliable attribution is essential. To allow flexible and adaptable designs, we introduce Constitution-Guided Watermarking, a framework that selects request-appropriate trade-offs from provider requirements, listed as natural-language principles. Offline, a pretrained reasoning agent examines constitutional rules alongside watermark implementations and iteratively refines rule-specific configurations using empirical feedback. At deployment, a separate monitor identifies applicable rules and retrieves the corresponding policy, including watermarking exemptions, without modifying the serving model. Furthermore, our framework supports offline parallel optimization and refinement of rule-specific configurations based on evolving provider requirements without affecting deployment, and binds each deployed configuration to its evaluation evidence, making deployment decisions auditable. In a proof-of-concept evaluation using KGW and a five-rule constitution, our framework selects configurations responsive to provider priorities and improves post-paraphrase detection on robustness-prioritized requests by up to percentage points over fixed configurations, while matching or exceeding all baselines in aggregate quality and clean detection at a nominal false-positive rate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.