acceptodds
Under review as a conference paper at ICLR 2027

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

Abstract

As deepfake detection models increasingly produce natural-language explanations, their reasoning often remains weakly grounded in visual artifacts, limiting reliability and user trust. Existing benchmarks mainly evaluate classification accuracy, overlooking whether explanations reflect the actual manipulations. This gap hinders progress toward deployable, explainable deepfake detection systems. To this end, we introduce XPlainVerse, a large-scale benchmark designed for joint deepfake detection and human-centered explanation. XPlainVerse comprises one million real and manipulated images, pairing authentic images from five established sources with forgeries generated by twelve off-the-shelf image editing and synthesis models. We further propose a multi-stage filtering pipeline, Edit-Check, to verify whether manipulations satisfy their intended edits, yielding high-quality, edit-consistent examples for reasoning supervision at scale. Beyond dataset scale, XPlainVerse provides two complementary explanation styles: technical explanations for expert analysis and simplified explanations designed for non-technical users. To evaluate explanation quality beyond surface similarity, we propose novel metrics, EntityScore and EvidenceScore, that assess whether explanations capture the manipulated entities and supporting evidence reflected in the reference explanations. Human annotations on 2,000 manipulated images, together with complementary human evaluation, validate explanation and edit quality. We believe XPlainVerse will make grounded explanation quality an explicit evaluation target in deepfake detection and support scalable research on trustworthy, interpretable models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.