A Living In-the-Wild Benchmark for Measuring the Static-Benchmark Gap in AI-Generated Video Detection
Abstract
Video detectors that succeed on laboratory benchmarks can fail on videos shared online, where generators, editing, and platform processing change over time. We introduce VidTide, a benchmark that tracks this changing population through successive, fixed releases. Its April 2026 slice contains 21,504 clips with provenance-based labels; a blind audit finds 96.4% agreement and 0.6% clear label errors. Five detector families score 52–71% AUROC on a balanced evaluation subset, compared with their published 93–96% on GenVideo. Standard backbones trained on VidTide reach 90.5% mean AUROC with linear probes and 91.7% when selecting the better of probing and fine-tuning. Controlled evaluations examine compression, metadata shortcuts, and newer multimodal models; temporal and cross-dataset tests assess transfer. We release versioned manifests and download scripts, and measure URL availability to document reconstruction limits. The results support updating training-data coverage while preserving fixed test sets. Code and data: https://anonymous.4open.science/r/vidtide-anon-C0B8/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.