SelfCheckMark: Confidence-Driven Dual-Channel Watermarking and Self-Correcting Extraction for Video Diffusion Models
Abstract
As video diffusion models gain wider use and generated videos become more prevalent, reliable copyright attribution and provenance tracing become increasingly important. In-generation watermarking addresses this need by embedding information during synthesis. A prominent route modifies the initial diffusion noise and recovers the embedded watermark information through video inversion. Yet generation, video distortions, and inversion errors yield recovered latent evidence of uneven reliability, so weak observations can dilute strong ones. Moreover, correcting an inverted noise estimate before combining redundant observations can fail when its errors exceed the code's tolerance. To address these challenges, we propose SelfCheckMark, a blind video watermarking framework with confidence-aware extraction and post-fusion self-correction. During embedding, extended Hamming(8,4) coding and a dual-channel design assign each coded bit to a strong-response anchor and continuous redundant elements. During extraction, local values are summed so elements with larger absolute values contribute more, while channel–frame replicas are weighted by inverse standard deviation, reducing the influence of unreliable evidence. Only after replica fusion is the most likely legal Hamming codeword selected, allowing correction to act on the combined evidence. Experiments show that, with a 512-bit watermark, SelfCheckMark achieves 100% mean bit accuracy on clean videos, over 98% under Gaussian noise and blur, and 85.36% under H.264 compression. Ablations confirm both confidence fusion and post-fusion correction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.