HumanForge: A Human-Centric Deepfake Video Benchmark with Multi-Agent Forgery Rationales

Authors: Wenbo Xu, Zhimin Chen, Xiaojie Liang, Hengrui Liu, Wei Lu

Published: 2026-07-09 17:12:41+00:00

Comment: 6 pages, 2 figures

AI Summary

The paper introduces HumanForge, a large-scale, human-centric deepfake video benchmark designed to address limitations in existing datasets concerning human-object/human-human interactions and multi-modal alignment. To construct and annotate HumanForge, the authors propose Gen2Anno, a modular active multi-agent pipeline that uses contrastive reasoning to generate high-fidelity videos and comprehensive annotations, including binary decisions, artifact categories, and spatio-temporal localization. Benchmarks reveal significant challenges for state-of-the-art detectors and Large Multimodal Models in zero-shot generalization and fine-grained reasoning on this new dataset.

Abstract

Rapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, posing unprecedented challenges to digital content forensics. Existing benchmarks primarily focus on either face-swapping or global text-to-video synthesis, overlooking the crucial dimensions of human-object or human-human interactions and multi-modal alignment. To address these limitations, we introduce HumanForge, a unified, large-scale, and multi-paradigm human-centric video forgery dataset. To construct and annotate this dataset without labor-intensive manual labeling or hallucinated monolithic prompts, we propose Gen2Anno, a modular active multi-agent pipeline built on LangGraph. Gen2Anno coordinates six specialized agents-ranging from source profiling to MoE-based reference analysis and closed-loop forensic verification-to generate over 18K high-fidelity video segments and produce structured, contrastive omni-annotations containing binary decisions, fine-grained artifact categories, and spatio-temporal localization. Extensive benchmarks using state-of-the-art traditional detectors and Large Multimodal Models (LMMs) demonstrate the significant challenges of zero-shot generalization and fine-grained reasoning on HumanForge. Code and dataset will be publicly released.


Key findings
The HumanForge benchmark, with its diverse human-centric scenarios and complex interactions, poses significant challenges for current deepfake detection methods. State-of-the-art traditional detectors and Large Multimodal Models exhibit difficulties in zero-shot generalization and fine-grained reasoning on this dataset. The Gen2Anno framework effectively generates detailed, contrastive omni-annotations by leveraging generative provenance, enabling more reliable and context-aware forensic analysis.
Approach
The authors developed HumanForge, a dataset of over 18,000 synthetic videos covering four human-centric scenarios (Audio-Driven, Pose-Driven, Interaction, Semantic-Driven), generated using more than ten state-of-the-art models. They also created Gen2Anno, a multi-agent framework that uses contrastive verification between expected and actual states to automate fine-grained, provenance-aware deepfake annotation, providing binary classification, artifact categorization, and natural-language explanations.
Datasets
HumanForge (newly introduced), HDTF, DFD, FFIW, FF++, SHHQ, TikTok
Model(s)
Wan2.1, CogVideoX, LTX-Video, OmniWeaving, SkyReels, Kling, Veo, InfiniteTalk, One-to-All, Animate-X, UniAnimate-DiT, HuMo, InfinityStar, AnchorCrafter, Hunyuan
Author countries
China