Ihor O. Serhiienko, Volodymyr O. Artemchuk · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.2669.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Cross-silo deepfake speech detection must accommodate organizations that observe different attack profiles and channel conditions while avoiding routine centralization of raw speech. We study federated learning (FL) for a compact detector combining linear-frequency cepstral coefficients (LFCCs) with frozen WavLM representations over heterogeneous partitions derived from ASVspoof 2019 and 2021. An official-metadata audit of the historical seed-42 Client-C split reproduces the reported 2,572/3,655 (70.37%) bona-fide source overlap; a stricter unified-source audit against the complete A+B+C fit pool finds still greater dependence. We therefore repeat the principal audio evaluation with five source-disjoint splits whose holdouts are independently verified against the complete fit pool. Mean equal error rate (EER) is 3.836% for pooled centralized training, 4.622% for Federated Averaging (FedAvg), 4.095% for local C, 21.883% for local A, and 33.398% for local B. On the common Client-C holdout, A-only and B-only models transfer poorly, while FedAvg is on average 0.787 percentage points worse than pooled training and 0.527 points worse than local C; these comparisons do not establish native-domain benefit for A or B. Separately, a 20-seed controlled stress test examines integrity mechanisms under malicious participation. Increasing untargeted label-flip compromise produces progressively more negative benign--malicious update alignment and stronger aggregate-step attenuation; the signature persists across four classifier topologies and FedProx settings. Additional tests cover targeted label poisoning, a controlled trigger backdoor, zero-update free riding, and coordinated sign-flip model poisoning. We frame the contribution as an auditable experimental protocol and a bounded set of observations, not as a new detector, optimizer, attack, aggregator, or universal defense.
No comments yet — start the discussion below.