acceptodds
Under review as a conference paper at ICLR 2027

XLSR-EMA: A 3D Joint Forgery Manifold for Universal Speech Deepfake Detection

Abstract

Pre-trained self-supervised learning (SSL) models have achieved remarkable success in speech deepfake detection. However, existing methods typically rely on the final hidden layer or simple layer aggregation, limiting their ability to exploit the hierarchical information encoded across SSL representations. This limitation becomes more pronounced for speech generated by modern Audio Language Models (ALMs), whose artifacts span multiple acoustic and semantic levels. We propose XLSR-EMA, a universal speech deepfake detection framework that organizes the hidden representations of a pre-trained XLS-R model into a unified Time-Layer-Feature Joint Forgery Manifold. An Efficient Multi-scale Attention (EMA) module models interactions among temporal, feature, and layer dimensions, while a Dual-Purpose Perceptual Augmentation (DPPA) strategy improves channel robustness and sensitivity to subtle synthetic artifacts. Experiments on multiple benchmarks demonstrate the effectiveness and cross-domain generalization of the proposed framework. Anonymous code is https://anonymous.4open.science/r/XLSR-EMA-for-UADD-2B19available online.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.