Rabbia Mahum, Mohammad Shehab, Emad Abouel Nasr, Mohammed El-Meligy, Amna Sarwar, Sarang Shaikh · Array 2026 · 2026
DOI: 10.1016/j.array.2026.101184
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
: While speech spoofing detection techniques based on deep learning have performed well in the recent years, many deep learning-based speech spoofing detection methods based on convolutional neural networks (CNN), Recurrent Neural Networks (RNN), and Transformer based architectures have difficulty in effectively representing discriminative local acoustic features and long-range dependencies, which reduces their resistance to the sophisticated attacks of synthesized and voice converted speech. To overcome these disadvantages, we introduce a new detector for automatic speech spoofing detection, which is called Transformer Encoder and Ensemble Learning-based Detector (TEEL-Det). The proposed framework combines the ensemble learning, Mel-Frequency Cepstral Coefficients (MFCCs), and Transformer Encoder to improve the representation of features and contextual modeling. The ensemble learning module extracts the complementary discriminative features and the MFCCs give compact and informative acoustic speech representations from raw speech. The Transformer Encoder uses a self-attention mechanism to learn the characteristics of the speech that are of interest for accurate classification of bonafide and spoofed speech while maintaining long range dependencies. The proposed model was tested on a subset of the Logical Access (LA) category of the ASVspoof 2019 challenge, which includes synthesized speech and voice converted speech. Through experimental results, the effectiveness and robustness of TEEL-Det are clearly shown in speech spoofing detection, and its performance is better than that of other methods, with an Equal Error Rate (EER) of 0.003% for the LA evaluation set.
No comments yet — start the discussion below.