acceptodds
Under review as a conference paper at ICLR 2027

Beyond Inversion: Source-Anchored Forgery Attacks on Semantic Watermarks for Diffusion Models

Abstract

Semantic watermarking encodes provenance signals in the initial noise of diffusion models without modifying model weights, enabling source attribution that is robust to common image transformations. Yet, malicious semantic edits can preserve these signals, allowing manipulated content to retain attribution to the original generator or user. Inversion-based forgery attacks first estimate an initial noise latent in an attacker-controlled model and then regenerate the image under a target prompt. Watermark retention along this path, however, is sensitive to mismatch between the attacker and victim models. We propose an inversion-free, source-anchored forgery attack that bypasses this reconstruction step. Using the observed image's latent encoding as a fixed anchor, the attack integrates differences between velocity predictions at paired source and target states to perform prompt-guided semantic manipulation. It requires neither the source prompt nor queries to the victim generator or watermark verifier. Experiments across five victim generators and multiple semantic watermark methods demonstrate consistently higher watermark retention than inversion-based baselines. The attack also produces semantic forgeries that pass watermark verification under a defense designed to bind watermarks to image semantics. These results expose a gap between robust source attribution and protection against semantic forgery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.