CHORD: Learning Watermark Receivers from Unmarked Pairs
Abstract
Watermarks embedded during autoregressive audio generation must support source verification after the generated tokens have been rendered, transformed, and re-encoded. We address this mismatch by learning how the observation channel transforms watermark evidence. For a sender with evaluable likelihood ratios, the observed ratio is the conditional expectation of the source ratio under the unmarked source–observation law. This identity yields a receiver-learning objective and a representation-selection criterion using only paired unmarked data, without waveform reconstruction or key-specific marked training examples. CHORD (CHannel-Observed Response Dictionary) realizes this objective with ordered-candidate sampling, whose rank actions define a small dictionary of local responses shared by all keys and payloads. Learned from unmarked pairs, these responses support composite scoring for joint owner identification and message recovery. On twelve-second music sessions, replacing identity transport with learned responses increases joint owner–payload recovery through two codec compositions from 385 and 379 to 501 and 505 out of 512 sessions, primarily by improving attribution. In speech, unmarked prediction loss selects the receiver, and learned responses yield more gains than regressions under neural codecs, including a held-out operating point. A frozen dictionary transfers to new owner keys without refitting. Together, these results make response learning a practical route to attribution and decoding after audio transformation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.