Detecting Audio-Visual Deepfakes Outside the Generator Objective
Abstract
A generator optimized for an audio–visual synchronization objective may match that objective at an observed sample while retaining a different local articulatory response. We propose Outside the Objective (OtO)}, an authentic-only test of how acoustic features inferred by a visual-to-acoustic predictor change under local mouth-tangent queries. Our central construction maps declared generator objectives into the acoustic response space and tests the complementary response conditioned on the observed audio. A canonical detector-space cosine gives an exact local plane; declared differentiable losses contribute Riesz representatives. Conditional learning, objective projection, and probe selection use one consistent residual. In a controlled synchronization-weight sweep, the zero-order cosine detector falls from 88.9 to 66.9 AUROC while the projected first-order test remains between 90.2 and 90.7. Under a common cross-domain protocol,OtO reaches 92.0% AUROC on MMDF and 84.1% on HiFi-AVDF, with 3.6 and 4.8 points attributable to objective-direction removal beyond audio conditioning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.