Sync When It Counts: A Modality-Conditional Framework for Robust Audio-Visual Deepfake Detection
Abstract
Current audio-visual forgery detectors primarily exploit cross-modal alignment, assuming a continuous synchronization between facial movements and the acoustic speech signal. While effective on curated benchmarks, this rigid paradigm catastrophically fails under real-world conditions. When a single modality is missing or degraded, state-of-the-art alignment models (e.g., AVH-Align) collapse to near-random performance. This reveals the vulnerability that enforcing cross-modal correlation effectively blinds the model to the rich, independent forensic signals within each surviving modality. To address this, we propose a modality-conditional, self-supervised framework that completely decouples unimodal representation learning from cross-modal synchronization. Our architecture pairs a contrastive audio-visual alignment module with independent visual and audio Joint-Embedding Predictive Architecture (JEPA) streams trained entirely without manipulation labels. During inference, a deterministic routing gate evaluates the input across two tiers. To prevent modality collapse, it routes inputs with missing modalities directly to the independent JEPA streams. For standard audio-visual inputs, it halts false-positive drift by explicitly dropping natural silences (e.g., authentic pauses or nods) from the synchronization calculation. Our method achieves state-of-the-art in-domain performance on AV-Deepfake1M++ (0.8525 AUC, 0.9434 AP). Crucially, the unimodal JEPA pathways retain robust detection capabilities under severe modality degradation (0.739 audio-only AUC and 0.673 visual-only AUC), whereas prior baselines fail entirely. Furthermore, in zero-shot cross-dataset evaluation on InDeepFake, our approach achieves 0.764 AUC, outperforming several supervised in-domain models. These findings demonstrate that substituting rigid alignment with conditional, context-aware routing is crucial for resilient deepfake detection in real-world scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.