acceptodds
Under review as a conference paper at ICLR 2027

Universal Representations for Robust AI-Generated Video Detection

Abstract

Methods for detecting AI-generated video fall into three long-standing categories: low-level analysis of generation artifacts, high-level tests of physical plausibility, and learned classifiers over deep embeddings. Detectors in each category generalize poorly beyond their training distribution, and, we argue, published comparisons between them are confounded by backbone size discrepancies. We present, to our knowledge, the first architecture that explicitly unifies all three categories of cues—density-based anomaly scores over frozen encoder features, fused with spectral and motion-coherence streams—and name it SEAM. Trained on a single generator from AIGVDBench, SEAM reaches 0.92 AUC on 19 unseen open-source generators, beating the state-of-the-art (SOTA) performance, and matching SOTA performance (0.87 AUC) on 11 unseen commercial generators, including Sora, Kling, and Pika. We further introduce SLOP, a dataset of 6,000 real and AI-generated videos from Instagram, TikTok, and X, and a taxonomy of such content. SEAM achieves the best in-domain performance on SLOP (0.96 AUC), and a SEAM trained on SLOP transfers to unseen commercial generators better than any baseline trained on the same data (0.89 AUC).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.