Teffic-Audio: Tell Fact from Fiction

Authors: Wan Lin, Li Wang, Jindong Wang, Kunyu Feng, Zhizheng Wu

Published: 2026-07-30 15:21:38+00:00

Comment: 16 pages, 1 figure, 7 tables. Technical report. Project page: https://tefficlabs.com/teffic-audio

AI Summary

Teffic-Audio is a general speech deepfake detection system that achieves robust generalization across heterogeneous spoofing mechanisms and audio conditions. Rather than complex architectures, it focuses on a strong training recipe involving multi-source data, balanced sampling, and diverse audio augmentation. Teffic-Audio significantly outperforms existing public systems on the Speech-DF-Arena leaderboard, setting a new benchmark for practical speech deepfake detection.

Abstract

Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and transmission channels. This variability makes robust generalization across heterogeneous conditions a central requirement for practical detection systems. This report presents Teffic-Audio, a general speech deepfake detection system designed for comprehensive evaluation environment. Teffic-Audio adopts a straightforward detector architecture consisting of a Conformer-based speech encoder, multi-head attentive statistics pooling, and a binary classifier. Rather than relying on additional architectural complexity, the system improves generalization through its training recipe, which integrates multi-source data, attack- and source-balanced sampling, and diverse audio augmentation. Trained only with open-source data, Teffic-Audio achieves a pooled EER of 1.454% on the 14 test sets of Speech-DF-Arena, outperforming all currently public systems on the leaderboard. It also obtains the lowest EER on five individual test sets and shows a favorable performance-complexity trade-off compared with larger leading systems. Overall, Teffic-Audio provides a strong and practical reference system for general speech deepfake detection.


Key findings
Teffic-Audio achieved a pooled EER of 1.454% on the 14 test sets of Speech-DF-Arena, outperforming all currently public systems. The training recipe, especially diverse audio augmentation, was found to be crucial for improving generalization, leading to substantial performance gains on challenging test sets. The study also highlighted the significant impact of encoder backbone, pooling layer, and encoder depth on overall system performance and the performance-complexity trade-off.
Approach
The system uses a straightforward Conformer-based speech encoder, multi-head attentive statistics pooling, and a binary classifier. Its main innovation lies in the training recipe, which leverages multi-source data, attack- and source-balanced sampling, and diverse audio augmentation to improve generalization.
Datasets
ASVspoof2015, ASVspoof2019LA, ASVspoof5, ADD2022, ADD2023 Track1, FakeOrReal, SpoofCeleb, ReplayDF, DFADD, MLAAD, LibriSeVoc, SpeechFake, Wavefake, CodecFake, LibriSpeech, AISHELL3, GigaSpeech, CNCeleb, CommonVoice
Model(s)
Conformer-based speech encoder (initialized from w2v-BERT 2.0 backbone), Multi-head Attentive Statistics Pooling (MHASP), MLP binary classifier
Author countries
UNKNOWN