Chain of Evidence: Localizing Video Forgeries via Image Forensic Evidence in Space and Time
Abstract
Diffusion-based video editing produces photorealistic forgeries that leave subtle forensic traces. These forgeries are often confined to small regions within a few frames, requiring precise localization across both space and time. Existing video forgery research has mainly focused on video- or frame-level detection, while joint spatial and temporal localization, which is required to verify footage used in news reporting or as legal evidence, has received less attention. Image forgery models provide spatial forensic evidence, but frame-wise inference cannot exploit temporal context to refine localization. We present Chain of Evidence, a parameter-efficient framework that extends perturbation-based diffusion forgery localization from images to video on a frozen SAM 3 backbone. The model learns spatial forensic evidence from images, and a lightweight 8.4M-parameter causal memory carries it across frames so that preceding context strengthens weak or ambiguous cues. Training on video updates only the causal memory and the presence head, leaving the image-trained forensic representation intact. We also propose a protocol that scores space and time jointly. Under it, chaining evidence across frames more than doubles spatiotemporal IoU over framewise inference on the out-of-distribution AINPAINT, from 0.162 to 0.420. On image benchmarks that no method was trained on, Chain of Evidence reaches an average localization IoU of 0.299, against 0.218 for the strongest baseline. On TVIL, it reaches a temporal F1 of 0.370, against 0.259 for the strongest baseline. Code is available at https://anonymous.4open.science/r/coe-iclr27.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.