acceptodds
Under review as a conference paper at ICLR 2027

Does AI-Generated Video Detection Really Need Temporal Modeling?

Abstract

Temporal information is a natural source of evidence for AI-generated video detection (AIGVD), yet under shortcut-controlled evaluation on GenVideo, a simple spatial-only detector reaches about 0.95 AUC, substantially outperforming existing temporal detectors at around 0.80. Rather than concluding that temporal modeling is unnecessary, we investigate why current methods fail to exploit temporal evidence effectively. Our analysis suggests a decomposition hypothesis: observed inter-frame variation entangles generation-related evidence with substantial non-generative variation from preprocessing, camera motion, and object motion. We examine two consequences of this entanglement. First, scalar summaries commonly used by current methods can map generation-related and non-generative changes to overlapping values, whereas element-wise differences preserve richer structure and substantially improve detection. Second, progressively reducing major non-generative components yields further consistent gains, making temporal evidence competitive with spatial evidence on GenVideo and stronger in AUC on the unseen GenVidBench benchmark. Moreover, the resulting debiased temporal evidence complements spatial evidence, with their fusion further improving performance. Across controlled representation analyses, progressive reduction of non-generative variation, and cross-benchmark evaluation, our results consistently show that properly modeled temporal evidence remains highly discriminative and valuable for AIGVD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.