TrajBind: In-Trajectory Content Binding for Forgery-Resistant Diffusion Watermarking
Abstract
Existing diffusion watermarking methods mainly consider distortion and removal attacks, while forgery attacks pose a different threat: an adversary may transfer valid watermark evidence to content that was not generated by the watermarking system. We introduce TrajBind, an in-trajectory, content-bound watermarking method that ties the watermark pattern to the generated content. At a late denoising step, TrajBind extracts a compact content code from the model's predicted clean image and combines it with a secret key to determine the watermark pattern embedded during the remaining denoising steps. This design requires no auxiliary generation or external semantic models, while verification only requires the received image. We further provide a general formulation of diffusion-watermark forgery that covers both detection-level and exact-message attacks, together with a theoretical analysis that bounds detection-level forgery success in terms of content-code collision. Across five existing forgery attacks, TrajBind has lower worst-case forgery success than each of the 12 baseline watermarking methods while maintaining robust watermark detection and message recovery under benign transformations and removal attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.