GenVidCLIP: Temporal-Aware CLIP with Multimodal Alignment for AI-Generated Video Detection
Abstract
The rapid advancement of AI-driven video generation has substantially increased the risk of highly realistic visual forgeries. Existing video forgery detectors often model video content in a homogeneous manner, making sparse and subtle forgery traces vulnerable to being overwhelmed by large amounts of normal spatial and temporal information, which leads to temporal feature dilution. To address this issue, we propose GenVidCLIP, a temporal-aware three-branch forgery detection framework built upon a Vision-Language Model. The core design explicitly decomposes video forgery representation into three complementary components: a Spatial Semantic Stream (SSS) for capturing global semantic context, a Temporal Kinematic Stream (TKS) for identifying sparse high-response temporal artifacts, and a text semantic branch for providing learnable authenticity-aware semantic priors. To effectively integrate these heterogeneous representations, we further introduce a Contrastive Residual Gating Network (CRGN), which models the divergence between spatial and temporal features and selectively injects localized temporal evidence into the global spatial representation before semantic alignment with text prototypes. Experiments on GenVidBench-6M show that GenVidCLIP achieves 85.05% ACC, 90.29% F1, 87.74% TNR on HD real videos, and 99.59% AUROC. On the cross-dataset GenVideo many-to-many protocol, it obtains 97.46% AP and 93.79% F1. Additional evaluations on Celeb-DF and AEGIS-Hard demonstrate transfer beyond the primary benchmark while revealing remaining challenges under diverse manipulation and generation distributions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.