Erase to Forge, Verify to Defend: Diffusion-based Forgery Attack and Self-verifiable Defense Mechanism
Abstract
Watermarking has emerged as a critical mechanism for establishing provenance of AI-generated media, yet its security against forgery attacks remains largely underexplored. Forgery attacks, wherein an adversary embeds a counterfeit watermark into arbitrary content to falsely attribute it to a target model, pose a severe threat to the integrity of watermarking systems. In this work, we introduce , a diffusion-based no-box forgery attack that repurposes watermark removal techniques to construct a proxy dataset of watermark-free images, from which a conditional diffusion model learns to estimate and reproduce the target watermark residual. By modeling the watermark as a structured residual rather than regenerating the full image, EraseForge achieves state-of-the-art forgery performance with as few as 1,000 training samples—a 5 improvement in data efficiency over existing methods. We further demonstrate that EraseForge circumvents the multi-message defense strategy previously proposed in the literature, achieving over 91% bit accuracy at samples where prior methods fail at . To address this threat, we propose a defense mechanism with two complementary variants—SeedBind and ContentBind—that transform verification from a finite message pool lookup into a structural consistency check, reducing forged bit accuracy to (near-random) across all tested watermarking schemes and attack strategies without modifying the underlying watermarking architecture. Specifically, SeedBind enforces deterministic consistency between a random seed and its pseudo-random expansion to neutralize estimator-based attacks, while ContentBind derives the watermark message directly from the image's learned visual features, additionally neutralizing residual-based one-shot forgery attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.