acceptodds
Under review as a conference paper at ICLR 2027

Joint Attention-semantic Disruption Attack on Non-autoregressive Text-to-speech Models

Abstract

Recent advancements in neural speech synthesis have shifted toward non-autoregressive models, but the adversarial robustness of such continuous-trajectory and non-token-based schemes remains under-explored. In this paper, we reveal that the mapping between the continuous acoustic prompt and the generation trajectory is highly susceptible to subtle cross-modal perturbations. We exploit this vulnerability and propose a novel black-box adversarial attack that simultaneously distorts acoustic embeddings, dissolves internal cross-modal alignment, and corrupts semantic representations. Extensive experiments show that the proposed attack successfully sabotages five state-of-the-art text-to-speech models with a small perturbation budget (SNRdB, audio waveform scaled to ) under eight representative adversarial defense measures: total variation regularization, Kullback-Leibler regularization, adversarial training, low-pass filtering, -law companding, randomized smoothing, the Griffin-Lim algorithm, and score-based speech enhancement. Our method can be applied to low-energy digital jamming, availability attack, and evaluation of generative audio robustness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.