SILO: Subspace-Preserving Spectral Injection for Generalizable Audio Deepfake Detection
Abstract
Audio deepfake detectors often struggle to generalize across unseen synthesis systems and real-world acoustic conditions. We first analyze the **complementary failure modes** of classical spectral features and self-supervised learning (SSL) representations: spectral cues expose transferable synthesis artifacts but are sensitive to acoustic variation, whereas SSL representations are more robust to such variation yet can miss fine-grained spoofing artifacts. Motivated by this observation, we propose **Spectral Injection via Low-rank Orthogonality (SILO)**, which uses utterance-adaptive spectral evidence to complement the robust SSL representation through a dedicated low-rank spectral-injection branch. SILO further introduces a **geometric constraint** that separates low-rank directions from their signed amplitudes and a **preserved-subspace constraint** that protects selected high-energy pretrained directions. Both constraints are maintained through optimization on restricted Stiefel manifolds. Across cross-dataset evaluations, SILO outperforms state-of-the-art methods, achieving 1.02% EER on ASVspoof 2019, 1.56% on ASVspoof 2021, and 14.38% macro-EER on MLAAD. We further introduce **Voice-Disorder Deepfake (VDD)** and **ForVoice Deepfake (FVD)**, two evaluation datasets targeting pathological speech and long-form forensic recordings, respectively. These datasets extend evaluation to these underexplored conditions and reveal substantial remaining challenges in pathological-speech and long-form deepfake detection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.