Can AI-Generated Video Detectors Keep Up?
Abstract
Detecting AI-generated videos (AIGVs) has become increasingly important as powerful video generation models continue to emerge. However, existing benchmarks provide limited indications of how detectors will perform in practical deployment for two key reasons: (1) strong performance on internal test sets often fails to transfer to broader data distributions; and (2) commonly used one-to-many protocols and non-temporal splits do not conform to actual application. We introduce AIGV-Annual, a million-scale, temporally partitioned and extensible benchmark for AI-generated video detection, covering videos from 100 fake generators and 13 real sources. AIGV-Annual establishes a cumulative many-to-many protocol, in which detectors are trained on generators released before each temporal cutoff and evaluated on those released afterward, enabling the measurement of temporal decay and update gain under continual generator evolution. Using AIGV-Annual, we systematically evaluate 18 supervised baselines and 16 zero-shot configurations across multiple temporal cutoffs. Our experiments reveal substantial out-of-distribution degradation, temporal decay and varying update gains from newer training data. We further introduce a spatial-temporal detector. Extensive experiments demonstrate that our method outperforms state-of-the-art methods on existing benchmarks and in-the-wild data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.