Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil

Authors: Lucas Rafael Stefanel Gris, Daniel Casanova, Frederico Santos De Oliveira, Alef Iury Ferreira, Beatriz Almeida Felício, Raul César Reis Mata, Anderson da Silva Soares

Published: 2026-07-30 18:44:22+00:00

AI Summary

This research introduces ParlaSpoof-BR, a new audio deepfake dataset of Brazilian Portuguese political speech, to benchmark state-of-the-art deepfake detectors and analyze biases. The study reveals that current detectors struggle with generalization to this domain, with methodological factors (e.g., synthesis model, manipulation extent) being more influential than demographic factors on detection performance. ParlaSpoof-BR serves as a crucial benchmark for developing robust deepfake detection systems for electoral integrity.

Abstract

Recent advances in generative artificial intelligence have made it easier to fabricate statements and amplify political disinformation during elections. We introduce ParlaSpoof-BR, an audio deepfake dataset derived from recordings of the Brazilian Chamber of Deputies and expanded with synthetic utterances from diverse text-to-speech and voice conversion models. Using ParlaSpoof-BR, we benchmark state-of-the-art audio deepfake detectors, examine their ability to generalize to Brazilian Portuguese political speech, and investigate potential biases in their predictions. Our analysis reveals that current systems struggle to provide consistent decisions across the diversity represented in the dataset, with methodological factors (synthesis model choice, manipulation extent) dominating over demographic disparities. ParlaSpoof-BR provides a domain-specific benchmark for studying audio deepfake detection in a socially consequential and underrepresented setting, supporting the development of more robust detection systems for electoral integrity in Brazil.


Key findings
State-of-the-art detectors show significantly degraded performance on ParlaSpoof-BR compared to traditional benchmarks, with high false-positive rates on genuine audio. Methodological factors, such as the specific synthesis model and the extent of partial manipulation, exhibit a far greater impact on detection evasion than demographic factors like gender or region. Notably, shorter partial manipulations are harder to detect, and certain lossy codecs (like OGG) can severely compromise detection accuracy on genuine audio.
Approach
The authors created ParlaSpoof-BR by expanding Brazilian parliamentary speech recordings with synthetic utterances from various text-to-speech and voice conversion models, including partial manipulations. They then benchmarked three state-of-the-art audio deepfake detectors (AASIST, AASIST-L, DF-Arena-1B) on this dataset. Their analysis focused on evaluating generalization to Brazilian Portuguese political speech and investigating biases related to synthesis methods, manipulation extent, and demographics.
Datasets
ParlaSpoof-BR (newly introduced), ASVspoof 2019 Logical Access (for detector training/baselines).
Model(s)
AASIST, AASIST-L, DF-Arena-1B.
Author countries
Brazil