Prompt-Decoupled AI-Generated Multimodal Video Detection and Attribution
Abstract
As AI-generated videos become increasingly realistic, deepfake video detection is vital for safeguarding cybersecurity and video source attribution helps ensure regulatory accountability. Existing deepfake detection methods focus on audio and vision modalities (e.g., the mismatch between lip movements and audio) while largely ignoring the textual modality. Meanwhile, research on video attribution is limited, especially from a multimodal perspective. To address these gaps, we propose a Prompt-Decoupled MultiModal Forensics (PDMF) that combines vision and text modalities for concurrent AI-generated video detection and attribution. PDMF converts video frames into descriptive text subtitles, extracts visual and textual features through advanced encoders, and projects textual features into visual space. By applying the cross-modal residual subtraction, we filter out shared semantic context and successfully expose model-specific artifact features. These features are stored in an extensible reference database to convert video attributes into instance retrieval tasks. By using only 100 videos per class and 2-10 times less time than alternative methods, PDMF achieves a binary deepfake detection accuracy of 96.3% (outperforming SOTA by over 14.1%) and an attribution mAP of 91.0% (surpassing competitors by over 29.3%) on GenVidBench. Additionally, PDMF demonstrates robust generalization under few-shot conditions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.