Predicting Token-Level Watermark Retention from Unwatermarked Images for Autoregressive Image Models
Abstract
Token-level watermarks make autoregressive (AR) image models more likely to sample tokens selected by a secret key, creating a signal that identifies watermarked images. However, the signal can remain detectable after image processing on one AR image model and become difficult to detect on another. Existing evaluation requires generating watermarked images for each model and watermark. In this paper we propose a metric for the signal retained when generated tokens are rendered into an image and recovered for detection. We prove that higher gives higher detection success for a fixed model, watermark, and sampling bias at a fixed false-positive rate under stated statistical assumptions. We then propose RenderGap, an evaluation method that predicts through a derived approximation using only unwatermarked generations. RenderGap compares generated tokens with tokens recovered from the images and scores both using the watermark key. Experiments with independent calibration on various AR image models show median absolute relative prediction errors of – across JPEG, blur, and tokenizer variants. RenderGap therefore evaluates retention for a given AR image model and watermark before watermarked images are generated.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.