Visual Thought Matters: Learning to Ground and Refine Continuous Reasoning for Generalizable Deepfake Detection
Abstract
Generalizable deepfake detection requires reasoning over subtle visual evidence that is difficult to faithfully express in discrete language. While multimodal large language models (MLLMs) can generate detailed reasoning chains, text-dominated reasoning may overlook fine-grained forensic cues. To this end, we propose COVE, a framework that integrates continuous forensic states into the reasoning trajectory to capture and leverage fine-grained visual evidence throughout reasoning. Complementary frozen visual experts supervise these states to encode appearance, semantic, and boundary cues, equipping the MLLM with fine-grained forensic perception. To encourage the model to use this evidence for decision making, we introduce a Counterfactual Token Utility Reward (CTUR), which quantifies the contribution of forensic states to the correct verdict through counterfactual suppression and optimizes their decision utility via reinforcement learning. We further introduce a contrastive grounding objective that contrasts the same target reasoning under semantically matched visual conditions and optimizes its grounding gain relative to a frozen reference model. This encourages the refined reasoning to depend on its supporting visual evidence, while complete-chain supervision and replay preserve reliable reasoning behavior. At inference, COVE uses a single MLLM to generate the complete forensic reasoning trajectory and authenticity verdict. Experiments on HydraFake demonstrate strong generalization across diverse out-of-distribution scenarios, particularly on unseen forgeries and data domains, while producing high-quality forensic reasoning. Together, these results demonstrate the effectiveness of COVE in connecting fine-grained visual representation, decision utility, and evidence-grounded reasoning for generalizable deepfake detection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.