acceptodds
Under review as a conference paper at ICLR 2027

Freeze the Weights, Not the Graph: Prompt-Retrievable Watermarks for Frozen Diffusion Models with Certified Retrieval

Abstract

Prompt-retrievable watermarking asks a text-to-image model to embed a semantic feature of the prompt in the generated image so that a verifier holding a bank of candidate prompts can later return the one that produced it. We realise this on a frozen Stable Diffusion 1.5 with a trainable token selector, latent injector and ResNet decoder around the fixed sampler, and show that two implementation defaults decide whether it works. The memory-cheap freeze, a no-grad denoise with a detached output latent, removes the only gradient into the injector and pins retrieval at chance, whereas a differentiable frozen-weight rollout with per-step gradient checkpointing keeps the weights fixed but the graph intact and lifts held-out Top-1 to 13 to 87 times chance across bank sizes. The latent clamp inherited from bounded-template watermarks saturates a third of a standard-normal latent and replaces the image with a colour field; clamp-free inference restores image content at unchanged retrieval, 0.357±0.025 Top-1 at , and turns the injection amplitude into a dial between retrieval and prompt fidelity. We state a randomized-smoothing certificate for retrieval with an explicit finite-sample ceiling and certify the model with votes over three seeds. A control suite built around the unwatermarked image from the same seed establishes that the decoder reads the watermark rather than the image, that the signal survives the training attack suite, and that near-duplicate prompts and abstention remain weaknesses. Code and checkpoints are released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.