Benchmarks Lie, Signals Reverse: A Forensic Audit of AI-Generated Video Detection
Abstract
AI-generated video detection has run an implicit measurement program for a decade: assemble real and generated video, train a separator, report benchmark accuracy. We audit the program and report its failure mode. Without precedent, to our knowledge, the field's strongest prior, motion amount, reverses sign across generator generations, fake-to-real flow ratio 0.362 on a 2024 generator against 3.342 on Seedance-2.5, and the 2026 wall is a detector-reference pairing: on Seedance every confounded-corpus system fails or inverts, and the standardized corpus's own system reads 0.4478 below chance on one reference family, above on another. The infrastructure failure is published, VidAudit's: of ten surveyed families four fall to zero-feature container rules that decode no frames, and on GenVidBench a duration threshold reaches a deployable train-fit 79.87 official while a constant FAKE prediction scores 80.00, above ten of eleven published systems. The surveyed signals fail our falsification battery: eight gates, each with a demonstrated kill, leave 0 of the 17 tested frozen-estimator families standing as a universal detector; the battery and family list co-evolved, so the zero reads as an exploratory falsification ledger, not a closed census. The systems layer survives only under a container-standardized protocol: 0.865 pooled AUC zero-shot on the one standardized benchmark, a hundred in-family labels reading 0.996 on the wall pool, near-duplicates unexcluded, off-family content ranking 0.744 with 8% detected. The limits are stated, not softened: single-seed rows, a small wild track, a frozen-estimator survey. The contribution is the protocol, the falsifications, and a label-priced adaptation ladder.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.