Defending Faces without Touching Pixels: Audio-Only Protection against Personalized Talking Face Generation
Abstract
The rapid development of talking face generation (TFG) has raised growing concerns about identity misuse. In particular, audio-driven 3D TFG can reconstruct a reusable personalized 3D portrait from a reference video and animate it with arbitrary audio, making identity impersonation increasingly accessible. To counter such unauthorized reuse, existing proactive defenses predominantly inject protective perturbations into visual content to disrupt generative identity acquisition. However, these defenses face a sharp tradeoff between defense effectiveness and perceptual quality. Perturbing facial regions can noticeably degrade portrait appearance due to the high perceptual sensitivity of human faces, whereas shifting perturbations toward less visually sensitive regions or the background makes them more vulnerable to resizing, compression, and background replacement. To mitigate these limitations, we propose VoxGuard, a perceptually constrained audio-only defense that disrupts the learning of correspondences between audio and identity-specific facial motion during personalization. Instead of perturbing visual content, VoxGuard modifies only the audio track of the reference video, leaving visual frames unchanged and thereby avoiding modification of visually sensitive facial regions. To further reduce perceptual distortion, we employ psychoacoustic masking to guide perturbations toward perceptually masked regions. We further optimize across heterogeneous audio encoders to enhance black-box transferability and introduce continual optimization to mitigate interference among encoder objectives while preserving previously established representation shifts. Extensive experiments demonstrate that VoxGuard effectively disrupts unauthorized 3D TFG while limiting perceptual degradation, exhibits black-box transferability, and remains effective after common audio transformations, including MP3 and AAC compression and audio resampling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.