MoCouple: Learning Fine-Grained Motion Coupling for Generalizable Talking-Face Forgery Detection
Abstract
Advances in talking-face generation (TFG) enable increasingly convincing audiovisual impersonation, heightening the risks of identity misuse and privacy violations. Visual artifact analysis and audiovisual synchronization analysis constitute two dominant paradigms in current detection methods, with the latter primarily focusing on speech and lip dynamics. However, generator-specific artifacts can limit transfer to unseen synthesis methods, while accurate lip synchronization can render audio–lip mismatch less discriminative. Lip-centered analysis can also overlook subtle inconsistencies across facial regions and become vulnerable when non-frontal views obscure mouth dynamics. Although recent TFG methods increasingly capture audio–motion correspondences, aligning each motion component with audio alone does not ensure natural coordination across motions. Motivated by this, we propose MoCouple to detect subtle cross-component inconsistencies by learning the joint coordination of lip articulation, facial deformation, head movement, and speech from genuine videos. Specifically, masked latent reconstruction on real videos learns how lip articulation varies jointly with speech, facial deformation, and head motion, while window-level motion statistics capture weaker dependencies between audio and motion. Fusion guided by quality and evidence then combines these relation sensitive features with visual forensic cues to detect forgeries with plausible local motions but inconsistent joint dynamics. Experimental results demonstrate that, in addition to maintaining strong performance on the original training dataset, our method achieves significant out-of-distribution generalization capabilities that surpass existing methods. This work positions fine-grained motion coupling as a complementary authenticity criterion beyond visual fidelity and audio–lip alignment. The same perspective may inform the evaluation and training of TFG models, encouraging coherent joint dynamics rather than merely plausible movements of individual regions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.