Learning Privacy-Preserving Representations for Audio-Visual Deepfake Detection
Abstract
We introduce Privacy-Preserving Audio-Visual Deepfake Detection (PPDD), which aims to suppress private information while preserving robust and generalizable deepfake detection. Existing privacy-preserving approaches primarily target unimodal tasks and are inadequate for audio-visual deepfake detection, where subtle spatial and temporal manipulation cues may coexist with privacy information, requiring fine-grained privacy suppression. To address this, we propose an end-to-end adversarial framework that learns privacy-preserving audio and visual representations using modality-specific anonymizers. A Feature Suppression Block (FSB) performs fine-grained feature-level suppression, guided by unimodal and cross-modal privacy supervision. Experiments on FakeAVCeleb demonstrate substantial privacy suppression with negligible loss in detection utility. The learned representations further preserve detection utility across unseen manipulations and datasets, transfer across detectors and modalities, suppress unseen privacy attributes, and remain resistant to inversion attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.