acceptodds
Under review as a conference paper at ICLR 2027

Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction

Abstract

Modern TTS systems increasingly generate synthetic speech at scale for diverse users. This setting calls for content-level provenance that can verify the origin of released speech and attribute it to the requesting user, which generative watermarking can support by embedding multi-bit identifiers directly into synthesized speech. Once released, however, speech may undergo heterogeneous learned transformations during distribution and editing, with reconstruction objectives that can preserve speech utility while affecting watermark recoverability differently. We find that no single watermark carrier remains consistently reliable across reconstruction models, as its survival depends jointly on the embedded structure, reconstruction mechanism, and observation representation. To this end, we propose Thrive, a multi-bit generative speech watermarking framework for modern autoregressive TTS, covering both discrete-token and continuous-representation generation under reconstruction attacks. Specifically, Rise synchronizes watermark injection into intermediate representations with its continued integration into subsequent generation, while Care combines waveform and spectral experts using bit-wise reliability selection. Experiments on both autoregressive paradigms show that Thrive preserves synthesis fidelity, achieves 87.6% average recovery accuracy under reconstruction attacks, and supports source attribution over candidate sets of up to 10,000 identities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.