CounterFake: Towards Generalizable MLLM-based Deepfake Detection
Abstract
Deepfake detection remains challenging under distribution shifts caused by unseen data domains and evolving forgery techniques. Multimodal large language model (MLLM)-based detectors show promising generalization but remain vulnerable to these shifts. We introduce CounterFake, an MLLM-based deepfake detector designed to improve robustness to various domain and forgery shifts. To improve detection of fake images from unseen domains, we develop two complementary counterfactual preference losses: an image-level loss using hard-negative images and a text-level loss using counterfactual rationales. These losses are supported by ReForensics (Reference-guided Forensics Annotation), a comparative captioning pipeline that uses reference images to generate forensic rationales grounded in decisive forgery cues. We further highlight two simple yet important design choices for MLLM-based deepfake detection: (i) weight-space interpolation improves robustness on real images from unseen domains, while (ii) intermediate representation probing reads out authenticity information directly from intermediate language model features while retaining the ability to generate forensic rationales. Extensive experiments on DF40-based benchmarks demonstrate that CounterFake outperforms existing state-of-the-art MLLM-based deepfake detectors, e.g., by 9.1 and 7.9 percentage points in average macro-F1 under cross-domain shifts and joint domain/forgery shifts, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.