acceptodds
Under review as a conference paper at ICLR 2027

SAEMark: Architecture-Agnostic Image Watermarking via Sparse Autoencoder Representations

Abstract

Watermarking diffusion models is essential for tracing the provenance and enabling the trustworthy authentication of generated content. However, existing methods typically rely on architecture-specific generation or inversion procedures, limiting their applicability to emerging generative paradigms such as flow matching. We observe that Sparse Autoencoders (SAEs) can independently learn internal representations for different generative architectures, yielding structured and semantically rich feature spaces that provide a new substrate for decoupling watermarking from the underlying generation process. Based on this insight, we propose SAEMark, the first SAE-based watermarking framework for diffusion models, which directly embeds and extracts watermarks in the SAE feature space, thereby decoupling watermarking from the specific generation process. This representation-level design enables a unified watermarking mechanism across heterogeneous generative architectures while substantially improving robustness against watermark removal and image manipulation attacks. We conduct experiments across multiple architectures, covering 8 representative watermarking methods and 9 watermark removal attacks. The results show that SAEMark achieves leading watermark robustness while preserving high generation quality, demonstrating the potential of SAE representations as a unified substrate for robust and architecture-agnostic watermarking across generative architectures. Code is available at https://anonymous.4open.science/r/SAEMark-exp-C3E0/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.