Integrity Diagnostics for Cross-Silo Federated Deepfake Speech Detection Under Heterogeneity and Poisoning

Cross-silo deepfake speech detection must accommodate organizations that observe different attack profiles and channel conditions while avoiding routine centralization of raw speech. We study federated learning (FL) for a compact detector combining linear-frequency cepstral coefficients (LFCCs) with frozen WavLM representations over heterogeneous partitions derived from ASVspoof 2019 and 2021. An official-metadata audit of the historical seed-42 Client-C split reproduces the reported 2572/3655 (70.37%) bona fide source overlap; a stricter unified-source audit against the complete A+B+C fit pool finds still greater dependence. We therefore repeat the principal audio evaluation with five source-disjoint splits whose holdouts are independently verified against the complete fit pool. Mean equal error rate (EER) is 3.836% for pooled centralized training, 4.622% for Federated Averaging (FedAvg), 4.095% for local C, 21.883% for local A, and 33.398% for local B. On the common Client-C holdout, A-only and B-only models transfer poorly, while FedAvg is on average 0.787 percentage points worse than pooled training and 0.527 points worse than local C; these comparisons do not establish native-domain benefit for A or B. Separately, a 20-seed controlled synthetic-embedding stress test examines integrity mechanisms under malicious participation. Increasing untargeted label-flip compromise produces progressively more negative benign–malicious update alignment and stronger aggregate-step attenuation; the signature persists across four classifier topologies and FedProx settings. Additional tests cover targeted label poisoning, a controlled trigger backdoor, zero-update free riding, and coordinated sign-flip model poisoning. We frame the contribution as an auditable experimental protocol and a bounded set of observations, not as a new detector, optimizer, attack, aggregator, or universal defense.

Authors

Institutions

Publication Details

Journal
Future Internet
Published
2026-10-09
DOI
https://doi.org/10.3390/fi18100544
Primary Topic
Privacy-Preserving Technologies in Data
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Integrity Diagnostics for Cross-Silo Federated Deepfake Speech Detection Under Heterogeneity and Poisoning

Volodymyr O. Artemchuk, Ihor Serhiienko
Future Internet
Privacy-Preserving Technologies in Data
article

Integrity Diagnostics for Cross-Silo Federated Deepfake Speech Detection Under Heterogeneity and Poisoning

Volodymyr O. Artemchuk, Ihor Serhiienko
article en

Abstract

Cross-silo deepfake speech detection must accommodate organizations that observe different attack profiles and channel conditions while avoiding routine centralization of raw speech. We study federated learning (FL) for a compact detector combining linear-frequency cepstral coefficients (LFCCs) with frozen WavLM representations over heterogeneous partitions derived from ASVspoof 2019 and 2021. An official-metadata audit of the historical seed-42 Client-C split reproduces the reported 2572/3655 (70.37%) bona fide source overlap; a stricter unified-source audit against the complete A+B+C fit pool finds still greater dependence. We therefore repeat the principal audio evaluation with five source-disjoint splits whose holdouts are independently verified against the complete fit pool. Mean equal error rate (EER) is 3.836% for pooled centralized training, 4.622% for Federated Averaging (FedAvg), 4.095% for local C, 21.883% for local A, and 33.398% for local B. On the common Client-C holdout, A-only and B-only models transfer poorly, while FedAvg is on average 0.787 percentage points worse than pooled training and 0.527 points worse than local C; these comparisons do not establish native-domain benefit for A or B. Separately, a 20-seed controlled synthetic-embedding stress test examines integrity mechanisms under malicious participation. Increasing untargeted label-flip compromise produces progressively more negative benign–malicious update alignment and stronger aggregate-step attenuation; the signature persists across four classifier topologies and FedProx settings. Additional tests cover targeted label poisoning, a controlled trigger backdoor, zero-update free riding, and coordinated sign-flip model poisoning. We frame the contribution as an auditable experimental protocol and a bounded set of observations, not as a new detector, optimizer, attack, aggregator, or universal defense.

Future InternetVol. 18(10)
Pukhov Institute for Modelling in Energy Engineering (UA), Palladin Institute of Biochemistry (UA)
Openalex Percentile: Top 13%
Privacy-Preserving Technologies in Data
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Integrity Diagnostics for Cross-Silo Federated Deepfake Speech Detection Under Heterogeneity and Poisoning — Volodymyr O. Artemchuk, Ihor Serhiienko · Future Internet (2026) | TGRS Research Map | TGRS