acceptodds
Under review as a conference paper at ICLR 2027

Persistent Watermarking of Text-to-Image Models

Abstract

Text-to-image (T2I) generation is gaining increasing popularity with the general public, motivating the development of reliable mechanisms for copyrighting such models given their expensive training costs. An adversary may obtain and reuse a pretrained T2I model without authorization, and then serve a modified version through an API service. Such modifications may arise from ordinary downstream adaptation or deliberate attempts to erase ownership, including input-prompt preprocessing, model fine-tuning, and output post-processing. From the model owner's perspective, a key challenge is therefore to embed trigger data that remain persistent under such changes while preserving the model's normal image-generation capabilities. In this work, we propose a contrastive-style watermarking objective with a term that explicitly encourages the watermarked model to behave *differently* from the original model on trigger inputs. Experiments show substantially stronger trigger-data persistence than prior methods across a wide range of common downstream modifications and deliberate attempts to weaken the watermark, resulting in substantially higher detection rates, often approaching 100% TPR@FPR<.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.