Periodic Gradient-Variance Suppression on SVD-Adapted Residuals for Cross-Generator Deepfake Detection
Abstract
As deepfake generation methods diversify rapidly, detectors trained on a fixed set of manipulation techniques increasingly fail to generalize to unseen generation paradigms, undermining their real-world reliability. Parameter-efficient adaptation of frozen vision backbones, such as SVD-decomposed low-rank residuals, enables efficient deepfake detection but offers no mechanism to distinguish generalizable forgery signal from manipulation-specific shortcuts within the few directions it is allowed to adapt. We study whether environment-conditioned gradient disagreement across known manipulation methods can identify which trainable spectral coefficients in a frozen CLIP residual are reliable, and whether this signal differs meaningfully from temporal gradient instability, the criterion used by concurrent work. Through a controlled, same-seed comparison, we find that suppressing high-disagreement directions fails when applied once, but periodically refreshing the suppression mask matches a Fishr-style cross-environment baseline across three paired seeds. At matched intervention strength, this criterion also outperforms random and magnitude-based suppression controls, supporting an effect beyond simple capacity reduction. A matched FMSD-style temporal control performs consistently below the cross-manipulation approaches in the evaluated seeds, suggesting the environment signal contributes information beyond gradient instability alone. Leave-one-manipulation-out (LOMO) evaluation reveals a substantial transfer asymmetry that is not explained by simple gradient similarity between methods, with one manipulation method's detection collapsing to near-chance performance when held out entirely. On four unseen generators spanning diffusion and GAN-based synthesis, our approach improves over baselines concentrated on the diffusion-based DDIM generator, while the diffusion-transformer case remains difficult for all methods tested.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.