CoSignature: Cross-Modality Watermarking with Verifiable Binding
Abstract
The rapid deployment of multimodal models has enabled the seamless generation of interleaved text and images. However, existing watermarking schemes treat these modalities in isolation, leaving systems vulnerable to misinformation campaigns where malicious actors manipulate context or swap single-modality content. To address this, we propose CoSignature, a cross-modality watermarking scheme that establishes a verifiable binding between visual and textual outputs via a shared content ID, which we hope to embed into both modalities. To overcome the limited capacity of text watermarking, we employ a novel “zero-bit" binding mechanism: rather than embedding the content ID as a payload, we use it to seed the text generation algorithm itself, converting binding verification into a statistical detection task that disambiguates the true content ID from noisy candidates retrieved by the image decoder. We provide theoretical guarantee that our binding error probability decays exponentially with text length. Empirical evaluation on four distinct architectures—native unified models (Janus-Pro-7B), composite pipelines (Qwen3-8B + Qwen-Image), Qwen3-VL-8B, and BAGEL-7B-MoT—demonstrates the efficacy of our approach, achieving binding accuracy on clean data, accuracy in detecting swap attacks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.