DiffMotion: Differential Facial Motion Modeling for Robust Deepfake Detection
Abstract
Recent advances in deepfake generation have made forged facial videos increasingly realistic, challenging detectors that rely on appearance artifacts. We propose, a landmark-sequence framework that models facial motion as a complementary forgery cue. DiffMotion combines raw coordinates with bidirectional lag-1 and lag-2 differences, propagates them over a muscle-inspired sparse facial graph, and aggregates local temporal context using folded, overlapping windows. This design provides an interpretable account of which facial regions and motion relations contribute to a decision. Experiments on FaceForensics++ and Celeb-DF show competitive within-domain performance and strong resilience to controlled compression and landmark perturbations; they also reveal a substantial cross-dataset gap. Notably, DiffMotion maintains an accuracy of over 86.75% under severe noise and high frame-drop rates, significantly outperforming existing methods.We therefore position DiffMotion as an anatomically motivated and degradation-resilient detector, while treating domain generalization as an important limitation rather than a solved problem.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.