CROSSMARK: Weaving Cross-Granularity Watermarks for Robust Detection of LLM-Generated Text and Code
Abstract
Watermarks for large language model (LLM) defined at a single granularity struggle to simultaneously achieve dense detection evidence, low generation distortion, robustness to local insertion, and resilience to rewriting. We introduce CrossMark, a cross-granularity watermarking framework that combines complementary evidence across token, local-unit, and global-unit levels. Its key idea is to let different granularities specialize in different attack regimes while reinforcing one another during detection. Dense token-level evidence supports reliable clean detection, local-unit keyed preferences expose insertion-induced inconsistencies, and rewrite-stable semantic representations preserve global evidence after substantial paraphrasing or code transformation. CrossMark further couples these signals rather than treating them independently: local consistency dynamically calibrates token statistics and identifies suspicious units for semantic verification, producing a unified joint detection score. Experiments on natural-language and code benchmarks show that CrossMark delivers the strongest or near-strongest clean detection, leads evaluated baselines under local-unit insertion and single-pass LLM rewriting, and preserves competitive text quality and code correctness. Ablations further reveal a clear division of robustness: local evidence is the primary contributor against insertion, whereas global semantic evidence provides the dominant signal under rewriting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.