acceptodds
Under review as a conference paper at ICLR 2027

One Bit, One Time Slot: Temporally Structured Watermarking for Audio Generation

Abstract

Audio generation models have made high-quality synthetic speech and music easy to produce, and equally easy to copy and redistribute without attribution, raising serious concerns over accountability and copyright. Watermarking offers a practical solution by embedding recoverable payloads into generated audio for ownership verification. Yet, existing efforts either apply the watermark post hoc, leaving it extrinsic to generation and removable by neural denoising or resynthesis, or inject it during generation as an unstructured global perturbation of the latent that entangles all bits and collapses recovery toward chance as the payload grows. In this work, we present TempoMark, a temporally structured generative watermarking framework that treats the generated audio as an ordered trajectory and binds each payload bit to its own temporal footprint. Specifically, TempoMark assigns each bit a dedicated time slot, expands it into chips with a spread-spectrum pattern, and injects the payload as a chip-aligned residual at an intermediate denoising step. Encoding and decoding share this bit-to-slot decomposition, under which each bit is recovered by matched filtering over its assigned chips. On top of this structured backbone, a lightweight neural residual in the decoder is trained to support payload readout, together with a small injection adapter as an acoustic-fidelity correction. Extensive experiments show that TempoMark outperforms state-of-the-art baselines in bit recovery accuracy and imperceptibility, and remains robust against audio perturbations and neural remapping attacks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.