acceptodds
Under review as a conference paper at ICLR 2027

Physics-Guided Cascaded Refocusing for Efficient Deepfake Video Detection

Abstract

Advances in generative models have dramatically improved the realism of manipulated face videos, posing escalating threats to public security, information credibility, and digital forensics. Dominant detectors follow a fixed-sampling, one-shot paradigm that overlooks the non-uniform temporal distribution of forgery evidence, where traces concentrate within only a few frames, rendering results unstable and subject to considerable randomness. A second active sampling stage promises gains in accuracy and stability yet confronts an intrinsic trade-off: if locating complementary evidence is costly, enlarging the initial sample would be cheaper. We propose a physics-guided cascaded forensic perception framework resolving this trade-off. In the first stage, a Local–Global Resonance Mechanism aggregates cross-frame CLS tokens into global forgery evidence and re-injects it into local patch representations, while predictive uncertainty lets high-confidence samples terminate early and others trigger re-inspection. Candidate clips are then ranked by feature discrepancy and filtered by abnormal facial motion, a discriminative yet network-free signal, enabling targeted acquisition of complementary evidence at negligible cost. Extensive experiments on multiple benchmarks, unseen manipulations, and cross-dataset settings demonstrate consistent improvements in accuracy, stability, and generalization with reasonable efficiency, while qualitative analyses provide interpretable evidence by highlighting facial regions and abnormal motion patterns.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.