Learning Reconstruction Errors as Spatiotemporal Evidence for AI-Generated Video Detection
Abstract
The proliferation of AI-generated videos has raised growing concerns about misinformation, fraud, and potential societal risks, creating an urgent need for detection. Reconstruction-error-based methods have shown strong empirical performance for AI-generated image detection by exploiting discrepancies between an input and its reconstructed version. Their extension to video detection, however, faces two challenges: (i) a fixed reconstructor is not optimized for the detection objective and may yield weak discriminative signals; and (ii) an additional raw-video feature stream is often introduced to model temporal dependencies, increasing architectural overhead. To address these challenges, we develop a detector that learns and exploits reconstruction errors as a unified source of spatiotemporal evidence. First, we jointly optimize a reconstructor and the downstream classifier to produce reconstruction errors that are more informative for detecting AI-generated videos. Second, we characterize temporal information using relative second-order temporal differences of the reconstruction errors. We further show that a 3D convolution over three consecutive errors can represent this relative second-order feature without explicitly constructing it, while simultaneously integrating spatial evidence. Extensive experiments on over 30 existing AI-video generators demonstrate competitive performance against existing training-based detectors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.