XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

Authors: Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh, Muhammad Haris Khan, Jianfei Cai, Abhinav Dhall

Published: 2026-07-03 18:52:42+00:00

AI Summary

This paper introduces XPlainVerse, a million-scale benchmark for explainable deepfake detection, featuring a million real and manipulated images with human-centered explanations. It addresses the gap in existing benchmarks by providing two explanation styles (technical and simplified) and novel metrics (EntityScore and EvidenceScore) to evaluate explanation quality beyond surface similarity. The benchmark aims to foster research into trustworthy and interpretable deepfake detection models.

Abstract

As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limiting reliability and user trust. Existing benchmarks mainly evaluate classification accuracy, overlooking whether explanations reflect the actual manipulations. This gap hinders progress toward deployable, explainable deepfake detection systems. To this end, we introduce XPlainVerse, a large-scale benchmark designed for joint deepfake detection and human-centered explanation. XPlainVerse comprises one million real and manipulated images, pairing authentic images from five established sources with forgeries generated by twelve off-the-shelf image editing and synthesis models. We further propose a multi-stage filtering pipeline, Edit-Check, to verify if manipulations satisfy their intended edits, enabling reliable reasoning supervision at scale. Beyond dataset scale, XPlainVerse provides two complementary explanation styles: technical explanations for expert analysis and simplified explanations optimized for non-technical users. To evaluate explanation quality beyond surface similarity, we propose novel metrics, EntityScore and EvidenceScore, that measure reasoning fidelity by checking whether explanations correctly identify manipulated entities and visual evidence. Human annotations on 2,000 explanation pairs validate our dataset quality against human judgment. We believe XPlainVerse will establish grounded explanation quality as a measurable dimension of deepfake detection and support scalable research on trustworthy, interpretable models.


Key findings
The XPlainVerse dataset successfully filters for high-quality, realistic manipulations and provides diverse, grounded explanations. Human evaluations showed that filtered fake images were difficult to distinguish from real ones, and dual-level explanations address different user needs effectively. Models fine-tuned on the dataset perform well in-distribution but generalize weakly to out-of-distribution data, indicating they learn generator-specific patterns rather than robust visual reasoning, highlighting the challenge for future research.
Approach
XPlainVerse creates its large-scale benchmark by pairing authentic images with forgeries generated by various image editing models, using a multi-stage filtering pipeline called Edit-Check to ensure manipulations satisfy intended edits. It offers both technical and simplified natural language explanations for manipulated images, and authenticity explanations for real images. Novel metrics, EntityScore and EvidenceScore, are proposed to measure the fidelity of these explanations.
Datasets
XPlainVerse (novel dataset with images from OpenImages, EMOTIC, PIPA, PIC 2.0, PISC, MultiFakeVerse), FaceForensics++, DFDC, DFFD, MultiFakeVerse, SemiTruths, DD-VQA, FFA-VQA, FakeBench, REVEAL-Bench, FakeXplained, SIDA, Holmes-Set, FakeClue, TRACE.
Model(s)
GPT-4o-mini, FLUX.2-dev, HunyuanImage, LongCat-Image-Edit, Qwen-Image-Edit-2511, GPT-Image-1.5, Seedream-4.5, Wan 2.6, Nano Banana 2, Nano Banana Pro, Gemini 2.0 Flash, GPT-Image-1, ICEdit, Gemini-3-Flash, GPT-5-mini, GLM-4.6V, Qwen3.5-27B, DeepSeek-V3.2, GPT-OSS-120B, Gemini-2.5-Pro, Qwen3.5-4B, LLaVA-1.6-7B, Qwen3-VL-8B, InternVL-3.5-14B, BusterX++.
Author countries
Australia, United Arab Emirates