Latent Diffusion Model Watermarks can be Removed by VAR Regeneration
Abstract
Diffusion regeneration has emerged as a universal removal attack against invisible image watermarking, but it is not able to remove latent-domain watermarks, while the underlying reason remains unaddressed. In this work, we reveal that the reason lies in the natural preservation of generation–inversion symmetry in diffusion regeneration. Motivated by this, we propose visual autoregressive regeneration (VARR), a simple yet powerful universal watermark removal attack that breaks this symmetry. Specifically, during the VAR model’s multi-scale quantization, it keeps low-scale semantic tokens while replacing high-scale ones with their predictions using VAR's autoregressive transformer. Subsequently, the partially modified token sets are used to reconstruct the image with the watermark removed. Across eight representative latent-domain watermarking methods, VARR achieves an average removal success rate of 92.8% while maintaining high visual fidelity with an average FID of 28.27, significantly lower than state-of-the-art regeneration methods (FID across 40 to 75). Furthermore, we show that the VAR domain is more effective for watermark removal compared with the pixel and frequency domains, empirically supporting the effectiveness of VARR. Our findings expose a fundamental vulnerability of latent-domain watermarks and establish autoregressive regeneration as a new paradigm for evaluating watermark robustness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.