SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

Authors: Linxi Li, Yuncong Yu, Qianwei Guo, Liwei Jin, Yechen Wang, Carsten Maple

Published: 2026-07-06 09:19:03+00:00

Comment: 7 pages, 1 figures

AI Summary

This paper introduces SynSFX, a large-scale dataset featuring 43,374 audio clips for detecting deepfakes in sound effects. It aims to bridge the research gap in non-speech audio forensics, as existing speech-centric detectors show limited generalization to synthetic sound effects. The dataset enables the study of isolated sound-effect deepfakes generated by various text-to-audio models.

Abstract

While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial resources, but remain limited in scale and generation provenance for studying isolated sound-effect deepfakes. To support this direction, we present SynSFX, a large-scale corpus of 43374 clips (26452 synthetic, 16922 real) spanning 7 popular text-to-audio models.


Key findings
Speech-centric deepfake detectors (AASIST, RawNet2) perform poorly on synthetic sound effects due to lack of speech-specific cues, while general audio models like EAT-AASIST show relatively better, but still limited, zero-shot generalization. Fine-tuning models solely on sound effects leads to catastrophic forgetting of speech deepfake detection capabilities. Joint-domain training mitigates this forgetting but reveals a persistent bottleneck in generalizing to sound effects from unseen generative models, indicating overfitting to specific, known generator artifacts.
Approach
The authors created SynSFX, a large-scale corpus of synthetic and real sound effects generated by seven different text-to-audio models. They evaluated existing deepfake detectors (AASIST, RawNet2, EAT-AASIST) in zero-shot and fine-tuned settings. Their approach focuses on understanding generalization issues and the phenomenon of catastrophic forgetting when adapting speech-centric models to sound effects.
Datasets
SynSFX (newly introduced), AudioCaps, Clotho, ESC-50, TACoS, WavCaps, ASVspoof 2019 Logical Access (LA), UrbanSound8K
Model(s)
AASIST, RawNet2, EAT-AASIST
Author countries
United Kingdom, USA