Parallel Hypothesis Decisions for Vision-Language Deepfake Face Detection
Abstract
Vision-language models (VLMs) are increasingly asked to judge whether a face image is real or forged, and the default recipe is to let them reason step by step before answering. We argue that authenticity judgment is a perceptual decision that is poorly served by slow, generative reasoning. With the same 4-billion-parameter VLM, the same forensic evidence, and the same images, reasoning for about 998 tokens before answering is not only slow but also inaccurate: on all six external benchmarks it scores below the forensic evidence it was given. We introduce PDM-Face, a detector that answers a panel of eight forensic hypotheses (diffusion, GAN, face swap, real photograph, abnormal spectral energy, up-sampling peaks, and two expert-evidence checks) in parallel and without generating a single token. The image and a compact set of evidence tokens from a gated forensic expert layer are encoded once; the key-value and recurrent states of this shared prefix are then copied into eight branches that are answered in one batched forward pass, and a learned fusion turns the eight three-way answers into a calibrated fake probability. Across six external face benchmarks, PDM-Face outperforms generative reasoning in all 12 comparisons (paired 95% confidence lower bound at least +0.051 AUROC) while being 21 times faster per image. It also significantly outperforms four recent detectors in 19 of 23 benchmark comparisons, improves over its own expert layer on the hardest video benchmarks (+0.034 on FaceForensics++ and +0.021 on DFDC, full test sets), and its parallel hypotheses beat a single-question decision layer on GenImage and 3rdBench for every seed and for a second VLM backbone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.