Revisiting Vocal-Tract-Based Detection for Modern Audio Deepfake Models
Abstract
Vocal-tract-based audio deepfake detection relies on the premise that deepfake generators may fail to reproduce the vocal-tract characteristics of natural human speech. However, its robustness against modern synthesis models and adversarial manipulation remains insufficiently understood. We systematically evaluate vocal-tract detection against recent text-to-speech (TTS) and voice conversion (VC) models and show that its effectiveness degrades substantially on modern models. Through ablation studies, we identify two major contributors to vocal-tract realism: learned linguistic representations in TTS models and explicit modeling of source-speech fundamental frequency in VC models. We further introduce VoicePatch, a model-agnostic post-generation editing attack that evades detection by replacing detector-sensitive phoneme segments with real speech. Across three evaluated deepfake models, VoicePatch increases detection bypass rates from below 2% to above 80%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.