acceptodds
Under review as a conference paper at ICLR 2027

How Robust Are Code Watermarks, Really? Rethinking Robustness Beyond Rule-Based Rewriting

Abstract

As code LLMs are increasingly used to generate, transform, and learn from software, code watermarks are becoming important for establishing output provenance and protecting training-data ownership. Their utility, however, depends on robustness to semantics-preserving edits. Existing evaluations primarily rely on predefined transformation suites, which characterize robustness only to hand-specified edits and leave open whether the same conclusions hold under learned, input-adaptive semantic rewriting. Existing alternatives only partially address this gap: prompted rewriters offer limited control over rewrite magnitude and functional preservation, while token-focused adaptive rewriters cover only a narrow subset of code transformations. We therefore introduce WARP-Washer, a key- and detector-agnostic learned rewriting stress test trained exclusively on non-watermarked code. WARP-Washer combines supervised rewrite learning with execution-grounded GRPO to learn broad, functionality-aware transformations. At inference time, it produces a single rewrite without access to the watermark key, detector, or execution feedback, allowing the same policy to stress-test heterogeneous watermarks without target-specific guidance. Across four public inference-time and dataset-level watermarks, four programming languages, and four model backbones, WARP-Washer substantially disrupts heterogeneous watermark carriers while maintaining high syntax validity or compilation success, as appropriate to each language. On SrcMarker, it reduces message recovery by 81–91% relative to predefined AST transformations while retaining over 98% test-pass rates. Under matched-fidelity evaluation, it also achieves higher valid-removal rates than prior adaptive rewriting approaches throughout the evaluated threshold range. These results show that robustness conclusions drawn from predefined transformations cannot be assumed to transfer to learned semantic rewriting, motivating learned rewriters as a complementary stress test for code-watermark robustness. Our code is available at https://anonymous.4open.science/r/WARP-Washer-626D.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.