acceptodds
Under review as a conference paper at ICLR 2027

Category-Oriented Universal Adversarial Perturbations for Deepfake Video Detectors

Abstract

Fake video detectors rely on different forensic evidence, including frame-level artifacts, pretrained visual representations, and temporal dynamics. This diversity makes it difficult for a universal adversarial perturbation to transfer across different detectors. We categorize existing detectors into three categories: Spatial-Artifact-Oriented (SAO), Semantic-Representation-Oriented (SRO), and Temporal-Dynamics-Oriented (TDO), and develop three category-specific attacks within a category-oriented framework. Specifically, we introduce frequency-aware optimization for SAO, a dual-band static perturbation for SRO, and frame-specific temporal perturbations for TDO. We further develop a unified attack that combines these complementary properties into a single structured perturbation and jointly optimizes it across the three detector categories. Experiments on four surrogate and six unseen detectors, multiple video generators, and external datasets show that the proposed ategory-specific attacks consistently improve transferability over conventional baselines. The unified attack further provides broader cross-category transfer, although with a trade-off in attack strength compared with separately optimized category-specific attacks. These results highlight the importance of detector forensic evidence in adversarial transfer and motivate robustness evaluation across different detection paradigms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.