Beyond Inversion: Source-Anchored Forgery Attacks on Semantic Watermarks for Diffusion Models
Abstract
Semantic watermarking encodes provenance signals in the initial noise of diffusion models without modifying model weights, enabling source attribution that is robust to common image transformations. Yet, malicious semantic edits can preserve these signals, allowing manipulated content to retain attribution to the original generator or user. Inversion-based forgery attacks first estimate an initial noise latent in an attacker-controlled model and then regenerate the image under a target prompt. Watermark retention along this path, however, is sensitive to mismatch between the attacker and victim models. We propose an inversion-free, source-anchored forgery attack that bypasses this reconstruction step. Using the observed image's latent encoding as a fixed anchor, the attack integrates differences between velocity predictions at paired source and target states to perform prompt-guided semantic manipulation. It requires neither the source prompt nor queries to the victim generator or watermark verifier. Experiments across five victim generators and multiple semantic watermark methods demonstrate consistently higher watermark retention than inversion-based baselines. The attack also produces semantic forgeries that pass watermark verification under a defense designed to bind watermarks to image semantics. These results expose a gap between robust source attribution and protection against semantic forgery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.