ByteMark: Multi-Bit Watermark for Bit-wise Image Generators
Abstract
Watermarking is a central tool for provenance tracing of AI-generated content, yet existing methods for bit-wise image generators are zero-bit: they can only distinguish between watermarked or clean (non-watermarked) data. Many real-world deployments demand more, such as tracing a generated image back to the specific model, user, or API key that produced it. We introduce ByteMark, the first multi-bit watermarking scheme for bit-wise image generators. Since bit-wise generators require an equal number of zeros and ones to achieve high fidelity, ByteMark encodes each message bit in the transitions within tokens, by inducing the next generated bit to be either repeated or flipped based on the value of the message bit. Because the rule is relative to the preceding bit rather than to a fixed value, it embeds the message without disrupting the bit distribution. We further extend this mechanism to payloads with more than a thousand bits by embedding different parts of a message into many partitions of a given image. We show that ByteMark preserves the visual fidelity of generated images, is highly robust to a wide range of watermark removal attacks, and that it exhibits radioactivity, i.e., new models trained on ByteMark-ed images inherit and produce the watermark. These properties enable ByteMark to provide fine-grained provenance tracing in bit-wise image generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.