Do people rely on ChatGPT more than their peers to detect deepfake news?

Authors: Yuhao Fu, Nobuyuki Hanaki

Published: 2026-08-02 23:24:44+00:00

AI Summary

This experimental study investigates human reliance on different sources (ChatGPT, human peers, linguistic experts) for detecting AI-generated deepfake news. Participants showed greater reliance on ChatGPT than human peers, and their performance improved when relying on high-quality advice, irrespective of its source. The findings emphasize that the effectiveness of AI-based detection tools depends on both their objective quality and public trust.

Abstract

This experimental study investigates how people rely on different sources of advice when detecting AI-generated fake news (deepfake news). In a laboratory deepfake detection task, student participants identified the proportion of human-written (non-AI-generated) content in synthetic deepfake news articles and received advice from ChatGPT (GPT-4), human peers, or linguistic experts. The results show that participants rely more on ChatGPT than on human peers when detecting GPT-2-generated deepfake news. Participants also rely more on linguistic experts than on peers, while the relative reliance on experts versus ChatGPT is mixed across experimental waves, potentially reflecting time trends in beliefs about AI-based detection. Importantly, in the additional experiment conducted in 2025 under the same experimental procedure, participants relied more on linguistic experts than on ChatGPT. Moreover, performance improvements reflect the joint role of reliance and advice quality, arising primarily when participants rely on high-quality advice. Overall, relying on AI to detect AI-generated deepfakes can improve detection outcomes, but only when AI-based detection tools are of sufficiently high quality. These findings highlight the dual role of GAI as both a source of deepfakes and a tool for mitigating related risks.


Key findings
Participants relied more on ChatGPT than on human peers for deepfake news detection. Performance improvements were driven by the quality of advice, with higher-quality advice leading to greater accuracy gains, regardless of the source. While early experiments showed higher reliance on ChatGPT than experts, later experiments under consistent conditions in 2025 indicated greater reliance on linguistic experts over ChatGPT.
Approach
The researchers conducted a laboratory experiment where student participants identified the proportion of human-written content in synthetic deepfake news articles. They then received advice from either ChatGPT (GPT-4), human peers, or linguistic experts before making a second identification. Reliance was quantified using the 'weight of advice' (WOA) metric, and performance was measured by accuracy and improvement.
Datasets
Japanese deepfake news collected from an open deepfake news dataset (primarily politics, sports, meteorology, public safety). The articles included totally real (human-written from Japanese Wikinews), totally fake (generated by OpenAI's Japanese GPT-2 model), and partially fake content.
Model(s)
ChatGPT (GPT-4) for generating advice, and OpenAI's Japanese GPT-2 model for generating deepfake news content.
Author countries
Japan, Cyprus