RLSpoofer: A Sample-Efficient Black-Box Spoofing Stress-Test for LLM Watermarks
Abstract
Large language model (LLM) watermarking has emerged as a promising approach for detecting and attributing AI-generated text, yet its robustness to black-box spoofing remains insufficiently evaluated. Existing evaluation methods often demand extensive datasets and white-box access to algorithmic internals, limiting their applicability. In this paper, we study watermark spoofing from a distributional perspective, viewing it as selectively reconstructing watermark-induced probability shifts under semantic constraints. We theoretically show that low-dimensional watermark-induced shifts can be learned more sample-efficiently than the full distribution, while KL-bounded reconstruction inherently favors larger watermark deficits. Guided by this analysis, we propose RLSpoofer, a reinforcement learning-based black-box spoofing attack that requires only 100 paired human–watermarked paraphrases and zero access to watermarking internals or detectors. Despite this weak supervision, it empowers a 4B model to achieve a 62.0% spoof success rate on PF-Watermark while preserving semantic fidelity, compared with below 7% for baselines trained on up to 10,000 samples. Our findings expose weaknesses in the spoofing resistance of current LLM watermarking paradigms and provide a sample-efficient framework for evaluating watermark robustness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.