Leveraging Redundancy in Representations for Robust Synthetic Video Detection
Abstract
Synthetic-video detectors are ranked on benchmarks whose default tier is raw or lightly compressed footage. However, when a video is streamed by a viewer, it is usually re-encoded by the distribution channel, often at a very different qual- ity. Across ten prior state-of-the-art detectors, we observe that the performance and ranking changes significantly based on the quality of the encoded video. For example, on the quality-controlled SynthForensics benchmark, the existing state- of-the-art detector falls from 93.2 at pristine quality to 44.6 %AUC at medium compression, which is indistinguishible from chance. By analyzing spatial fre- quency sub-bands, we find that the bands driving detector decisions shift with compression. Relying on specific bands–either by design or as a consequence of training–can therefore leave a detector vulnerable to compression. To reduce this dependence, we randomly suppress frequency bands during training so that no single band is indispensable and combine features from multiple network depths to provide alternative sources of evidence. We also match compression conditions between real and synthetic training samples, preventing compression itself from serving as a shortcut for identifying synthetic videos. Trained on a single generator, our detector holds 90.0 %AUC at CRF40 on the same benchmark, moving from 99.7 on the default tier. Beyond compression, we evaluate robustness to adversar- ial attacks, which remains understudied in synthetic-video detection. Applying band suppression at inference mitigates most of the accuracy loss caused by a query-limited black-box attack, without impacting pristine accuracy
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.