acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Image Watermarking in the Era of Frontier Generative Models

Abstract

Image watermarks are increasingly used to establish the provenance of AI-generated content. We demonstrate a watermark-removal attack that repeatedly regenerates a watermarked image with GPT-Image-2. The attack requires only API access, without knowledge of the watermarking method, detector access, or local model deployment. It achieves 89% average removal across 13 post-hoc and in-generation methods, with a removal-fidelity trade-off comparable to leading specialized attacks. The evaluated content-based watermarks resist regeneration. However, we show that an existing method that inserts visual concepts from a fixed private inventory is vulnerable to concept extraction. Our extraction attack recovers 79% of its concepts from 3,000 watermarked images. The recovered candidates enable 93.9% removal after ten targeted edits, at a cost to image fidelity, and 70.4% spoofing success when generating images outside the watermarking system. We propose K-SemMark-R, a registered content-based watermark that independently samples prompt-compatible carrier values for each image and stores verified carriers in a private registry. Experiments show that most initially detected images remain detectable by carrier verification after regeneration, although the method does not enroll every prompt. These findings highlight the potential of content-based watermarking and motivate research into image-specific designs that resist regeneration without relying on a shared concept inventory.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.