Detecting AI-Generated Video: A Vision-Language Dual-View Survey

Authors: Dylan Xinming Hou, Juntian Zhang, Xu Gu, Yichen Wu, Nils Lukas, Gus Xia, Xiuying Chen, Yuhan Liu

Published: 2026-07-12 14:25:04+00:00

Comment: 51 pages, accepted by ACL 2026

Journal Ref: Association for Computational Linguistics 2026 pages 32221 to 32255

AI Summary

This paper surveys AI-generated video (AIGC-V) detection, reframing it as Factual Fidelity Verification to assess consistency with real-world facts. It proposes a Vision-Language Dual-View taxonomy classifying existing methods across four layers: intrinsic cue analysis, spatiotemporal consistency, cross-modal consistency, and language-guided world-level reasoning, based on a review of 221 works.

Abstract

The evolving realism of AI-generated Videos (AIGC-V) is rapidly rendering traditional artifact-centric detection insufficient, necessitating a paradigm shift from low-level inspection to high-level semantic verification. This paper presents a comprehensive survey of AIGC-V detection, reframing the task as Factual Fidelity Verification, which asks whether the events, entities, and physical processes depicted in a video are consistent with real-world facts. To systematize this rapidly evolving field, we propose a Vision-Language Dual-View taxonomy that organizes existing methods into a hierarchical, four-layer landscape, spanning intrinsic cue analysis, spatiotemporal consistency modeling, cross-modal consistency reasoning, and language-guided world-level reasoning. This dual-view framing highlights a fundamental transition from artifact matching in traditional deepfake detection to evidence-based semantic verification enabled by vision-language models and agentic reasoning pipelines. Based on a systematic review of 221 works, we synthesize AIGC-V generation paradigms, survey the landscape of detection methods, and review evaluation metrics and benchmarks in line with proposed views. Finally, we discuss current challenges and identify promising directions toward robust, explainable, and trustworthy detection.


Key findings
Traditional artifact-centric detection is insufficient for increasingly realistic AI-generated videos, necessitating a shift towards high-level semantic verification and factual fidelity. The proposed Vision-Language Dual-View taxonomy effectively categorizes current methods and highlights a trend towards integrating vision-language models and agentic reasoning. The survey also identifies the critical need for robust, explainable, and trustworthy evaluation metrics and benchmarks that go beyond simple authenticity scores towards evidence-based justifications.
Approach
The paper surveys existing AIGC-V detection methods and categorizes them into a Vision-Language Dual-View, four-layer taxonomy. This framework progresses from low-level visual artifact detection to high-level semantic and factual verification, utilizing both visual and language-based cues. It details different methodologies within each layer, addressing the evolving nature of AI-generated content.
Datasets
FaceForensics++, Celeb-DF, DFDC, DeeperForensics-1.0, ForgeryNet, WildDeepfake, KoDF, CDDB, DF-Platter, DeepfakeBench, AI-Face, DD-VQA, ExDDV, FAQ / Beyond Static Artifacts, FakeAVCeleb, LAV-DF, FakeMix, AV-Deepfake1M, ArEnAV, MAVOS-DD, DigiFakeAV, SocialDF, VCapAV, X-AVFake, AV-Deepfake1M++, MMDF, GVF, GenVidDet, GenVideo, DVF, GenVidBench, Deepfake-Eval-2024, LOKI, GenBuster-200K, GenWorld, Ivy-Fake, DAVID-X, GenBuster++, DeeptraceReward, ER-FF++set, AEGIS, ViFBench, Video Reality Test, AIGVDBench, SynthForensics, MintVid, VideoPhy, Physics-IQ, IPV-Bench, Morpheus, T2VPhysBench, PhyWorldBench, VideoPhy-2, Physion-Eval, WorldSimBench, Towards World Simulator (PhyGen-Bench), StoryEval, T2VWorldBench, VideoVerse, SVBench, TRAVL, SPOTLIGHT, VideoHallu, PhyDetEx
Model(s)
3D CNN, Transformer, EfficientNet, ViT, CLIP adapter, LLaVA, Qwen2.5-VL, InternVL3, VideoLLaMA3
Author countries
United Arab Emirates, China, United States