acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Vision Foundation Models for High-Fidelity Forgery Detection by Diagnosing Optimization Collapse

Abstract

Vision Foundation Models (VFMs) offer a promising route to generalizable deepfake detection, but it remains unclear whether representations learned through general-purpose pretraining retain the subtle, transferable forensic cues needed to detect high-fidelity forgeries. We expose this representational weakness through Optimization Collapse, in which linear probes trained on frozen VFM features for high-fidelity forgery detection degrade to chance-level performance even under small adversarial parameter perturbations. We quantify this vulnerability using the perturbation radius at which collapse occurs, termed the Critical Optimization Radius (COR). Theoretical analysis, supported by empirical evidence, shows that the COR is positively associated with the Gradient Signal-to-Noise Ratio (GSNR), a measure linked to generalization, motivating the use of COR as a diagnostic for assessing the strength of transferable forensic cues retained in VFM representations. We validate that higher forgery fidelity exacerbates this optimization vulnerability in both cross-manipulation comparisons and controlled within-manipulation interpolation studies, as reflected in lower COR and poorer detector generalization. Layer-wise analysis further reveals the mechanism underlying Optimization Collapse: for high-fidelity forgeries, COR and GSNR progressively decline with depth, indicating that transferable forensic cues are increasingly attenuated as features propagate through the frozen backbone. Reducing the perturbation radius can avoid collapse but does not recover the attenuated forensic cues. To mitigate this collapse by improving GSNR, we propose the Contrastive Regional Injection Transformer (CoRIT), which uses feature differences as a forward-computable proxy for the gradient signal. CoRIT pools token features based on regional gradient-proxy alignment to reduce gradient-proxy variance, injects the pooled features into auxiliary tokens to preserve regional signals, and fuses intermediate and final-layer representations. CoRIT keeps the backbone frozen and optimizes only a lightweight classification head. Experiments show that CoRIT increases both COR and GSNR while achieving the highest average AUC among the compared methods in cross-manipulation and cross-dataset evaluations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.