SemAdapt: Adaptive Semantic Watermarking with Paraphrase Robustness for Large Language Models
Abstract
Watermarking large language model (LLM) outputs is a promising approach for establishing content provenance and mitigating misuse. However, LLM-generated text can be easily paraphrased by users or automated tools, disrupting token-level watermark statistics and reducing detection reliability. Although existing methods attempt to improve robustness against such paraphrases, they often do so without adequately considering text quality, limiting their practical applicability. We propose SemAdapt, an adaptive semantic watermarking method for robust LLM watermarking under paraphrasing. SemAdapt injects watermark signals at the level of semantic vocabulary buckets rather than individual tokens, and refines the vocabulary partition using lexical substitutions observed under paraphrasing. To balance detectability, paraphrase robustness, and text quality, SemAdapt further adapts watermark strength according to uncertainty over the same semantic bucket space, with controller parameters selected offline using multi-objective Bayesian optimization under noisy paraphrase evaluations. Extensive experiments demonstrate that SemAdapt achieves consistently stronger robustness against diverse paraphrase attacks than existing baselines, while maintaining comparable text quality and detectability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.