What a Deepfake Score Does Not Tell You: Same-Source Controls for Face-Forgery and Generated-Video Detectors
Abstract
Real/fake performance alone does not determine how a detector responds to non-generative processing, neural re-rendering, or cross-identity conditioning. We introduce source-matched controls for face-forgery and generated-video evaluation. For 136 FaceForensics++ test excerpts and two face-swap generators, we retain each original (R) and construct a pipeline-only arm (N), a self-swap (S), and a cross-identity swap (F) with the same implemented downstream pipeline. Across 29 released checkpoints from five families, several assign high forgery scores to pipeline-only videos (N/R AUC up to 0.97; at a common 0.5 threshold two video pipelines judge 63–65% of them fake against 1% of originals). The conditional discrimination between cross-identity and self-swap outputs (F/S) varies across the 26 primary checkpoints (0.47–0.79) and is weaker than the conventional forgery-versus-original discrimination for every one of them. On a third generator, GHOST, absent from the documented detector-training methods, N/R exceeds 0.5 for all 26 primary checkpoints and F/S is below F/R for 25. On FaceForensics++ test crops at threshold 0.5, Gaussian blur (σ=4) turns 99–100% of initially correctly classified real frames into false alarms for 16 of the 19 frame-level checkpoints, and JPEG quality 30 turns 15–96% of correctly detected NeuralTextures frames into misses for the 12 FaceForensics++-trained ones. In a controlled Xception system, pipeline-matched negatives suppress the positive pipeline response. With fixed SimSwap four-arm inputs and five epochs, changing only the self-swap target raises F/S from 0.55 to 0.97 on the training generator and from 0.56 to 0.69 on GHOST (not confirmed on InSwapper), and a shared-backbone two-head model supports both targets. A source-matched generated-video case study shows detector- and implementation-dependent responses to format conversion and same-family VAE reconstruction. Standard AUC does not resolve these operational sensitivities; we recommend reporting matched non-generative controls, conditional contrasts, and the inference and control-construction settings alongside detection results.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.