Randomized Representation Smoothing for Audio Watermarking
Abstract
With the rapid development of audio generators, audio watermarking has become an important technique for content authentication and provenance verification. In this work, we propose RRS-AW, a method for embedding watermarks into randomized transformation-averaged robust representations of audio during generation. The method encodes binary messages through pairwise coordinate ordering in representations averaged over random audio transformations and optimizes the generation latent while keeping both the generator and watermark extractor fixed. We provide a theoretical analysis of the robustness of the resulting representation and complement it with an empirical analysis of its local sensitivity. Experiments on AudioLDM-generated music demonstrate strong robustness across individual, composite, and codec transformations: RRS-AW achieves the best or tied-best bit accuracy on 9 of 13 composite attacks and 9 of 13 codec settings among the competitors. Additional experiments demonstrate robustness across different audio generators, low false-positive detection, and high perceptual audio quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.