Beyond Generator-Specific Shortcuts: Open-World AI-Generated Video Detection via Adversarial Training
Abstract
AI-generated video (AIGV) detection is increasingly important for mitigating misinformation, impersonation, and other malicious uses of synthetic media, yet existing detectors often generalize poorly to unseen generators in open-world settings. To better understand this challenge, we analyze the failure modes of AIGV detectors under cross-generator out-of-distribution (OOD) shifts and observe that detector representations strongly encode generator-specific cues, which are associated with substantial ID–OOD performance drops. Further analysis shows that naively perturbing generator-sensitive tokens can even reduce OOD performance, indicating that generator-related cues and detection-relevant evidence are only partially separable. Motivated by these findings, we propose Forensic Obfuscation of geneRator-specific cues via Gradient-guided Embedding perturbation (FORGE), a token-aware adversarial training framework for open-world AIGV detection. FORGE identifies tokens that are strongly associated with generator-specific shortcuts yet less critical to real/fake discrimination, and selectively perturbs them during training to suppress generator-dependent patterns while preserving useful forensic evidence. Through synchronized adversarial training between an attacker and a detector, FORGE encourages the detector to learn more transferable generator-invariant forensic representations. Experiments on 29 held-out benchmark entries show that FORGE achieves an average F1 of 92.62, outperforming eight strong baselines by 5.31 points, while also remaining robust under four common post-processing attacks. We will release our code and datasets upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.