EviPatrol: Evidence-Grounded Split-Modal MLLM for Audio-Visual Deepfake Localization and Diagnosis
Abstract
Audio-visual deepfake localization requires identifying the manipulated modality, precise temporal boundaries, and supporting evidence. However, existing methods often retain localization-related cues implicitly within latent representations or prediction scores, limiting their interpretability and downstream reasoning. Meanwhile, multimodal large language models (MLLMs) struggle to capture weak and localized boundary signals from raw audio-visual inputs due to their emphasis on semantic understanding. We propose EviPatrol, an explicit boundary evidence modeling framework for fine-grained audio-visual deepfake localization and diagnosis. EviPatrol transforms implicit localization cues into structured, time-indexed evidence representations through a Boundary Evidence Interface before MLLM reasoning. To improve evidence reliability, we introduce BLA-AMR and Safe-AMR to enhance modality-aware evidence learning and reduce cross-modal shortcuts. The resulting evidence is integrated with audio-visual content through a split-modal MLLM for independent audio and visual localization with evidence-grounded diagnosis. Extensive experiments on LAV-DF and AV-Deepfake1M demonstrate that EviPatrol achieves superior temporal localization and cross-dataset generalization. On LAV-DF, EviPatrol improves [email protected] by 51.94 percentage points over the previous best method, validating the effectiveness of explicit boundary evidence modeling under strict localization criteria.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.