DELTA: Detecting and Localizing Generated Video Content through Generative Latent Differences
Abstract
Partially manipulated videos pose a growing forensic challenge, as generated content can be inserted into short intervals while the rest of the video remains authentic. Existing approaches mainly derive temporal cues from frame-level discriminative representations. We argue that focusing on individual representation states can overlook how video content evolves over time. Based on this view, we introduce DELTA, a training-free, multi-scale descriptor that characterizes local transitions along generative latent trajectories. By modeling local evolution rather than isolated latent states, DELTA captures temporal transition structure beyond frame-level discriminative representations. For full-video detection, we map continuous spatial transitions into reusable discrete patterns for robustness. We further introduce a benchmark that evaluates temporal localization of partial forgeries using paired real sources and exact interval annotations. DELTA achieves strong performance across both temporal localization and full-video detection. In particular, on the unseen-generator evaluation, replacing CLIP features with DELTA under the same ActionFormer localizer improves mean AP from 77.68 to 86.31.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.