Alka Rani, Akshar Maitray, Anush G.R., Raut Anand · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23096915
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Audio deepfake detectors degrade on conditions unseen in training. Aggregation–separationdomain generalization (ASDG) was proposed to make detectors robust to an unseen generator;we ask whether the same mechanism transfers to a different axis, the language of the speech.We formalize language-as-domain ASDG, with real-only adversarial aggregation and a tripletseparation loss anchored on the real-speech centroid, and evaluate it on four Indian languagesfrom two families (Hindi and Marathi, Indo-Aryan; Kannada and Telugu, Dravidian), usinga CNN–Transformer classifier over frozen XLSR-53 features and a 51,075-clip corpus inwhich every fake is a paired neural re-synthesis of a real utterance (EnCodec, Vocos, DAC).Single-language baselines lose up to 3.9 EER points on a new language, but the loss doesnot follow language family (mean gap 0.0016 within family vs. 0.0039 across), and poolingall four languages is by far the strongest baseline (0.0325 EER). ASDG does not beat amatched empirical-risk-minimization (ERM) control (target-language EER 0.0728±0.0078 vs.0.0713±0.0044 for Hindi→Kannada; 0.0684±0.0058 vs. 0.0649±0.0046 for Hindi→Marathi),and a mechanism ablation over four transfer directions separates no variant from the others.When an unseen generator is added, ASDG is worse than ERM (0.330 vs. 0.284 and 0.402vs. 0.326 EER). Augmentation recovers 2–5 EER points under codec and noise degradation,including held-out types, while a contrastive term adds nothing measurable, and IntegratedGradients attributions are faithful and no less consistent across languages under ASDG. Wereport this as a negative result with its limits stated explicitly.
No comments yet — start the discussion below.