Ambiguity-Aware Temporal Transport for Audio-Visual Deepfake Detection
Abstract
Learning audio-visual consistency from authentic videos reduces dependence on specific manipulation patterns in deepfake detection. However, sustained or recurring visual states may produce similar matching responses across neighboring audio segments, leaving local correspondences ambiguous. Under such uncertainty, local matching strength may become less reliable as detection evidence. This motivates evaluating both correspondence informativeness and the temporal coherence of matching evidence. We propose Ambiguity-Aware Temporal Transport (AATT), a detection framework that jointly models these two properties. AATT estimates correspondence informativeness from the normalized entropy of local matching distributions to weight evidence for temporal inference. The weighted evidence is integrated across time through marginalization over monotone paths, yielding soft correspondences between the two modalities. These correspondences guide the aggregation of audio features onto the visual timeline for an informativeness-weighted comparison with visual features. The detection score combines the resulting transport compatibility and mean correspondence informativeness with path surprisal, which measures the cost of imposing temporal constraints. The framework learns exclusively from authentic talking videos through a ranking objective that distinguishes matched audio-visual pairs from cross-video mismatches. Experiments demonstrate substantial improvements on talking benchmarks, a pronounced advantage in the more challenging singing setting, and robustness to different perturbations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.