Faithful Is Not Forensic: Robust Intervention-based Forensic Testing (RIFT) for Auditing Deepfake Detectors
Abstract
A deepfake detector can be accurate and faithfully explained while relying on evidence that does not account for the manipulation that changed an image's authenticity. We introduce RIFT, Robust Intervention-based Forensic Testing, a post hoc audit for frozen detectors that uses only scalar scores. Given a matched pristine and manipulated pair, RIFT applies the same intervention to a candidate region in both images. Manipulation Reliance (M) measures how much that region accounts for the matched authenticity contrast; Nuisance Instability (Q) is a complementary fixed-family stability diagnostic; and FSS is a conservative joint summary. Across six admissible detectors, the matched effect exceeds score-matched same-authenticity placebos by 0.198 to 0.666, with every paired interval above zero. The audit also withholds interpretation where raw metrics would mislead: a detector that passes the in-domain competence gate and shows a positive raw GT-over-control margin fails the placebo test (−0.035, 95% CI [−0.050, −0.019]), so no forensic claim is made for it. A shared-cue sanity check and geometry-matched spatial controls confirm that the matched construction behaves as defined, while operator changes, reserved nuisances, a failed prespecified stress test, and controlled FF++ residual transfer bound the admissible interpretation. RIFT requires no retraining, gradients, internal features, or architecture-specific hooks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.