Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents
Abstract
Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Prior agent watermarks bind the signal to raw strings and break in two ways. AgentMark's own robustness test paraphrases the observation alone and bit-recovery falls to 16.8%; its released script holds the per-step key fixed, so that experiment measures distribution drift, not the key desynchronization its context-hash design separately admits. All three prior schemes also bind the watermark to the exact action symbol, so renaming a tool desynchronizes decoding even when the observation is untouched — the axis we measure and close. And every prior agent watermark studies only removal: none asks whether an adversary can forge a trajectory that verifies as someone else's, a question answered affirmatively for text watermarks (Jovanović et al., 2024). We present Semantic Behavioral Watermarking (SBW): watermarking over semantic action clusters under history conditioning, with the public-cluster bin replaced by keyed collision-resistant binning whose fresh-bucket assignment is provably unpredictable in the random-oracle model. Across five agent models (3B–14B, four vendors) and three encoders the ordering holds on both benchmarks: on ToolBench (600 trajectories per model) detection under rewriting is 0.49–0.66 for cluster-level versus 0.05–0.17 for exact-symbol at a permutation-calibrated 1% FPR, at 72–83% choice agreement against 22–27% for logit biasing; on ALFWorld (100 episodes per model) it is 0.92–0.97 versus 0.00–0.01. Keyed binning takes adaptive forgery from 100% to the false-positive floor at the primary operating point (bge, r=64). We also mark the boundary that guarantee does not cover: when the adversary copies the victim's own steps, shuffled splicing is neutralized (0.000 on Qwen2.5-3B) but chained replay remains at 0.76–0.98 across the five models, reported as open. Paraphrase robustness costs about half of the per-step watermark capacity. Code is available at https://anonymous.4open.science/r/SBW-Agent-Watermark.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.